Device, method, and graphical user interface for interacting with three-dimensional environment
The computer system enhances user interaction in augmented and virtual reality environments by using microgestures, hand-ready states, and line of sight detection to reduce input complexity and improve feedback, resulting in a more intuitive and efficient user experience.
Patent Information
- Application Number
- JP2025036597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-09-23
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-08
AI Technical Summary
Existing methods for interacting with augmented and virtual reality environments are cumbersome, inefficient, and place a significant cognitive burden on users, often requiring multiple inputs and providing insufficient feedback, leading to energy waste and suboptimal user experiences, particularly in battery-operated devices.
A computer system with improved methods and interfaces that utilize microgestures, hand-ready state configurations, line of sight detection, and hand grip changes to simplify interactions in three-dimensional environments, reducing the number and complexity of inputs and enhancing user feedback.
The system provides a more intuitive and efficient human-machine interface, reducing cognitive load and improving user safety and satisfaction by allowing natural, discrete interactions in three-dimensional environments.
Smart Images

Figure 2025102795000001_ABST
Abstract
Description
Technical Field
[0001] (Related Applications) This application claims priority to U.S. Provisional Patent Application No. 62 / 907,480, filed on September 27, 2019, and U.S. Patent Application No. 17 / 030,200, filed on September 23, 2020, and is a continuation of U.S. Patent Application No. 17 / 030,200, filed on September 23, 2020.
[0002] The present disclosure generally relates to a computer system having a display generation component, including but not limited to, an electronic device that provides virtual reality and mixed reality experiences via a display, and one or more input devices that provide a computer-generated experience.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include digital images, videos, text, icons, and virtual objects such as buttons and other graphics.
[0004] However, the ways and interfaces for interacting with environments (such as applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which the operation of virtual objects is complex and error-prone impose a significant cognitive burden on the user and degrade the experience in the virtual / augmented reality environment. In addition, those ways are time-consuming more than necessary, thereby wasting energy. The latter problem is particularly critical in battery-operated devices. SUMMARY OF THE INVENTION
[0005] Accordingly, there is a need for a computer system having improved methods and interfaces for providing a computer-generated experience that makes the interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can complement or replace conventional methods of providing a computer-generated reality experience to the user. Such methods and interfaces reduce the number, degree, and / or type of inputs from the user by assisting the user in understanding the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.
[0006] The above-mentioned deficiencies and other problems associated with the user interface for a computer system having a display generation component and one or more input devices are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generation component, and the output devices include one or more haptic output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or sets of instructions stored in the memory for performing a plurality of functions. In some embodiments, the user interacts with the GUI through a stylus and / or finger contact and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the user's body when captured by a camera and other motion sensors, and voice input when captured by one or more audio input devices.In some embodiments, the functions executed through the interaction optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game play, making a phone call, video conferencing, sending an email, instant messaging, training support, digital photography, digital video shooting, web browsing, playing digital music, taking notes, and / or playing digital video. The executable instructions for performing those functions are optionally included in a non-transitory computer-readable storage medium or other computer program products configured to be executed by one or more processors.
[0007] There is a need for an improved method and interface for interacting with a three-dimensional environment. Such a method and interface can complement or replace conventional methods for interacting with a three-dimensional environment. Such a method and interface reduce the number, degree, and / or type of inputs from a user and generate a more efficient human-machine interface.
[0008] In some embodiments, the method is executed in a computer system including a display generation component and one or more cameras, and the method includes displaying a view of the three-dimensional environment; while displaying the view of the three-dimensional environment, using one or more cameras to detect movement of a user's thumb on the user's index finger of a first hand; in response to detecting movement of the user's thumb on the user's index finger using one or more cameras, performing a first action according to a determination that the movement is a swipe of the thumb on the index finger of the first hand in a first direction; and performing a second action different from the first action according to a determination that the movement is a tap of the thumb on the index finger at a first position on the index finger of the first hand.
[0009] In some embodiments, the method is executed in a computing system that includes a display generation component and one or more input devices, and the method displaying a view of a three-dimensional environment; detecting a hand at a first position corresponding to a portion of the three-dimensional environment while the three-dimensional environment is being displayed; in response to detecting the hand at the first position corresponding to the portion of the three-dimensional environment, displaying a visual indication of a first operating context for gesture input using hand gestures within the three-dimensional environment according to a determination that the hand is maintained in a first predetermined configuration; and not displaying a visual indication of the first operating context for gesture input using hand gestures within the three-dimensional environment according to a determination that the hand is not maintained in the first predetermined configuration.
[0010] In some embodiments, the method is executed in a computer system that includes a display generation component and one or more input devices, and the method displaying a three-dimensional environment, including displaying a representation of a physical environment; detecting a gesture while the representation of the physical environment is being displayed; in response to detecting the gesture, displaying a system user interface within the three-dimensional environment according to a determination that the user's line of sight is directed towards a position corresponding to a predetermined physical position within the physical environment; and performing an operation in the current context of the three-dimensional environment without displaying the system user interface according to a determination that the user's line of sight is not directed towards a position corresponding to a predetermined physical position within the physical environment.
[0011] In some embodiments, the method is executed in an electronic device that includes a display generation component and one or more input devices, and the method Displaying a three-dimensional environment that includes one or more virtual objects, detecting a line of sight directed at a first object within the three-dimensional environment, the line of sight meeting a first criterion and the first object responding to at least one gesture input, and in response to detecting a line of sight that meets the first criterion and is directed at the first object in response to at least one gesture input, displaying an indication of one or more interaction options available for the first object within the three-dimensional environment according to a determination that the hand is in a predetermined ready state for providing a gesture input, and not displaying an indication of one or more interaction options available for the first object according to a determination that the hand is not in a predetermined ready state for providing a gesture input.
[0012] There is a need for an improved method and an electronic device with an interface that facilitate the use of a user of an electronic device interacting with a three-dimensional environment. Such a method and interface can complement or replace conventional methods for facilitating the user of a user of an electronic device interacting with a three-dimensional environment. Such a method and interface can create a more efficient human-machine interface, enable the user to further control the device, and the user can use a device with an improved safety, reduced cognitive load, and improved user experience.
[0013] In some embodiments, the method is executed on a computer system including a display generation component and one or more input devices, detecting an arrangement of the display generation component at a predetermined position relative to a user of the electronic device, and in response to detecting the arrangement of the display generation component at a predetermined position relative to a user of the computer system, displaying, through the display generation component, a first view of a three-dimensional environment including a pass-through portion, the pass-through portion including a representation of at least a part of the real world surrounding the user, detecting a change in the hand grip on a housing physically coupled to the display generation component while the first view of the three-dimensional environment including the pass-through portion is being displayed, and in response to detecting the change in the hand grip on a housing physically coupled to the display generation component, replacing the first view of the three-dimensional environment with a second view of the three-dimensional environment, the second view replacing at least a part of the pass-through portion with virtual content, according to a determination that the change in the hand grip on a housing physically coupled to the display generation component meets a first criterion.
[0014] In some embodiments, the method is executed on a computer system that includes a display generation component and one or more input devices, and includes displaying, via the display generation component, a view of a virtual environment; detecting a first movement of a user within a physical environment while the view of the virtual environment is being displayed and while the view of the virtual environment does not include a visual representation of a first portion of a first physical object that exists within the physical environment in which the user is located; in response to detecting the first movement of the user within the physical environment, changing the appearance of the view of the virtual environment in a first manner that indicates physical characteristics of the first portion of the first physical object, where the first physical object has a range that may be visible to the user based on the user's field of view with respect to the virtual environment, and in accordance with a determination that the user is within a threshold distance of the first portion of the first physical object, without changing the appearance of the view of the virtual environment to show a second portion of the first physical object that is part of the range of the first physical object that may be visible to the user based on the user's field of view with respect to the virtual environment; and preventing the appearance of the view of the virtual environment from being changed in the first manner that indicates physical characteristics of the first portion of the first physical object in accordance with a determination that the user is not within a threshold distance of the first physical object that surrounds the user within the physical environment.
[0015] According to some embodiments, a computer system includes a display generation component (e.g., a display, a projector, a head-mounted display, etc.), one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, one or more processors, and a memory storing one or more programs, the one or more programs being configured to be executed by the one or more processors and including instructions for performing or causing to be performed any of the operations of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium has instructions stored therein that, when executed by a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), and optionally one or more tactile output generators, cause the device to perform or cause to be performed any of the operations of the methods described herein. According to some embodiments, a graphical user interface of a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, a memory, and one or more processors for executing one or more programs stored in the memory includes one or more of the elements displayed in any of the methods described herein, and these elements are updated in response to input as described in any of the methods described herein. According to some embodiments, a computer system includes a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, and means for performing or causing to be performed any of the operations of the methods described herein.According to some embodiments, an information processing apparatus for use in a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch sensing surface, and optionally one or more sensors for detecting the intensity of contact with the touch sensing surface), and optionally one or more haptic output generators, includes means for performing, or causing to be performed, any of the operations of the methods described herein.
[0016] Accordingly, a computer system having a display generation component is provided with improved methods and interfaces that interact with a three-dimensional environment and simplify the use by a user of the computer system when interacting with the three-dimensional environment, thereby enhancing the effectiveness, efficiency, and user safety and satisfaction of such a computer system. Such methods and interfaces can complement or replace conventional methods for interacting with a three-dimensional environment and facilitating the use by a user of the computer system when interacting with the three-dimensional environment.
[0017] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The functions and advantages described herein are not exhaustive, and in particular, many additional functions and advantages will be apparent to those skilled in the art in view of the drawings, the specification, and the claims. Further, it should be noted that the language used herein has been selected solely for readability and for the purpose of explanation and not for the purpose of defining or limiting the subject matter of the invention.
Brief Description of the Drawings
[0018] To better understand the various embodiments described, reference should be made to the following "Modes for Carrying Out the Invention" in conjunction with the following drawings, and like reference numerals refer to corresponding parts throughout the following figures.
[0019]
Figure 1
[0020]
Figure 2
[0021]
Figure 3
[0022]
Figure 4
[0023]
Figure 5
[0024]
Figure 6
[0025]
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 7E
Figure 7F
Figure 7G
Figure 7H
Figure 7I
Figure 7J
[0026]
Figure 7K
Figure 7L
Figure 7M
Figure 7N
Figure 7O
Figure 7P
[0027]
Figure 8
[0028]
Figure 9
[0029]
Figure 10
[0030]
Figure 11
[0031]
Figure 12
[0032]
Figure 13
[0033] The present disclosure relates to a user interface for providing a user with a computer-generated reality (CGR) experience, according to some embodiments.
[0034] The systems, methods, and GUIs described herein improve user interface interaction with virtual / augmented reality environments in multiple ways.
[0035] In some embodiments, the computer system enables a user to interact with a three-dimensional environment (e.g., a virtual or augmented reality environment) using microgestures performed by small movements of a finger relative to other fingers or parts of the same hand. The microgestures are detected using a camera (e.g., a camera integrated with a head-mounted device or installed away from the user (e.g., in a CGR room)), as opposed to, for example, a touch-sensitive surface or other physical controller. Different movements and positions of the microgestures and various movement parameters are used to determine an action to be performed in the three-dimensional environment. Capturing the microgestures using the camera and interacting with the three-dimensional environment allows the user to freely move around the physical environment without being obstructed by physical input devices, thereby enabling the user to explore the three-dimensional environment more naturally and efficiently. Additionally, the microgestures are discrete, non-intrusive, and suitable for interactions that can be performed in public and / or require propriety.
[0036] In some embodiments, a hand-ready state configuration is defined. The additional requirement that the hand be detected at a position corresponding to a portion of the three-dimensional environment in which the hand is displayed ensures that the hand-ready state configuration is not misrecognized by the computer system. The hand-ready state configuration is used by the computer system as an indication that the user intends to interact with the computer system in a predetermined operational context that is different from the currently displayed operational context. For example, the predetermined operational context is one or more interactions with devices outside the currently displayed application (e.g., games, communication sessions, media playback sessions, navigation, etc.). The predetermined operational context optionally includes, for example, a home or start user interface that can start other experiences and / or applications, a multitasking user interface that can select and resume recently displayed experiences and / or applications, or a control user interface for adjusting one or more device parameters of the computer system (e.g., display brightness, audio volume, network connection, etc.). By using a special hand gesture to trigger the display of a visual indication of the predetermined operational context for gesture input that is different from the currently displayed operational context, the user can easily access the predetermined operational context without cluttering the three-dimensional environment with visual controls and without accidentally triggering an interaction in the wrong operational context.
[0037] In some embodiments, a physical object or a part thereof (e.g., a user's hand or a hardware device) is selected by the user or a computer system to be associated with a system user interface (e.g., a control user interface for a device) that is not currently displayed in a three-dimensional environment (e.g., a mixed reality environment). When the user's line of sight is directed to a position within the three-dimensional environment other than the position corresponding to a predetermined physical object or a part thereof, gestures performed by the user's hand cause operations to be executed in the currently displayed context without displaying the system user interface. When the user's line of sight is directed to a position within the three-dimensional environment corresponding to a predetermined physical object or an option thereof, gestures performed by the user's hand cause the system user interface to be displayed. Based on whether the user's line of sight is directed to a predetermined physical object (e.g., the user's hand performing the gesture or the physical object the user is attempting to control using the gesture), the user selectively executes operations in the currently displayed operation context or displays the system user interface in response to an input gesture, so that the user can efficiently interact with the three-dimensional environment in two or more contexts without visually cluttering the three-dimensional environment with multiple controls, and the interaction efficiency of the user interface is improved (e.g., the number of inputs required to achieve a desired result is reduced).
[0038] In some embodiments, the user's line of sight directed to a virtual object within a three-dimensional environment in response to a gesture input causes visual indications of one or more interaction options available for the virtual object to be displayed only if it is found that the user's hand is in a predetermined ready state to provide the gesture input. If the user's hand is not found to be in a ready state to provide the gesture input, the user's line of sight directed to the virtual object does not trigger the display of the visual indication. By using the combination of the user's line of sight and the ready state of the user's hand to determine whether to display a visual indication of whether the virtual object has associated interaction options for the gesture input, the user is not unnecessarily exposed to the constant changes when shifting the line of sight around the three-dimensional environment, and useful feedback is provided to the user when the user explores the three-dimensional environment with their own eyes, reducing the user's confusion during the exploration of the three-dimensional environment.
[0039] In some embodiments, when a user positions the display generation component of the computer system at a predetermined position relative to the user (e.g., placing a display in front of the eyes or placing a head-mounted device on the user's head), the view of the real-world user is blocked by the display generation component, and the content presented by the display generation component dominates the user's view. Sometimes the user benefits from a smoother and more controlled process for transitioning from the real world to a computer-generated experience. Thus, when displaying content to the user via the display generation component, the computer system displays a pass-through portion that includes at least a partial representation of the real world surrounding the user, and replaces at least a portion of the pass-through portion with virtual content for display only in response to detecting a change in the user's hand grip on the housing of the display generation component. The change in the user's hand grip is used as an indication that the user is more ready to transition to a more immersive experience than what is currently presented through the display generation component. The gradual transition of the immersive environment controlled by the change in the user's hand grip on the housing of the display generation component is intuitive and natural for the user and improves the user's experience and comfort when using the computer system for a computer-generated immersive experience.
[0040] In some embodiments, when a computer system displays a virtual three-dimensional environment, the computer system applies a visual change to a portion of the virtual environment at a position corresponding to the position of a portion of a physical object that enters within a threshold distance of the user and that can be within the user's field of view with respect to the virtual environment (e.g., that portion of the physical object becomes visible to the user, but the presence of the display generation component blocks the user's view of the real-world around the user). Further, instead of simply presenting all portions of physical objects that can be within the field of view, portions of physical objects that are not within the user's threshold distance are not visually presented to the user (e.g., by changing the appearance of portions of the virtual environment at positions corresponding to those portions of the physical objects that are not within the user's threshold distance). In some embodiments, according to the visual change applied to portions of the virtual environment, one or more physical characteristics of portions of physical objects that are within the user's threshold distance are presented within the virtual environment without completely stopping the display of those portions of the virtual environment or completely stopping the provision of an immersive virtual experience to the user. This technique enables warning the user about physical obstacles approaching the user as the user moves around within the physical environment while exploring an immersive virtual environment without unduly invading and disturbing the user's immersive virtual experience. Thus, a safer and smoother immersive virtual experience can be provided to the user.
[0041] Figures 1-6 illustrate an exemplary computer system for providing a CGR experience to a user. Figures 7A-7G show exemplary interactions with a three-dimensional environment using gesture input and / or gaze input, according to some embodiments. Figures 7K-7M show exemplary user interfaces that are displayed when a user transitions an interaction with a three-dimensional environment, according to some embodiments. Figures 7N-7P show exemplary user interfaces that are displayed when a user moves through a physical environment while interacting with a virtual environment, according to some embodiments. Figures 8-11 are flowcharts of methods for interacting with a three-dimensional environment, according to various embodiments. The user interfaces of Figures 7A-7G are used to illustrate the processes of Figures 8-11. Figure 12 is a flowchart of a method for facilitating use of a computer system by a user for interacting with a three-dimensional environment, according to various embodiments. The user interfaces of Figures 7K-7M are used to illustrate the processes of Figures 12-13.
[0042] In some embodiments, as shown in FIG. 1, the CGR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., household appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted device or a handheld device).
[0043] When describing the CGR experience, various related but distinct environments are individually referred to using various terms for the sake of the user to perceive and / or interact with (e.g., using inputs detected by the computer system 101 to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the CGR experience) the various inputs that the user can interact with. The following is a subset of these terms.
[0044] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the aid of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through senses such as vision, touch, hearing, taste, and smell.
[0045] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to an environment that is wholly or partially simulated, through which people can perceive and / or interact via an electronic system. In CGR, a subset of a person's body movements or their representation is tracked, and in response, one or more characteristics of one or more virtual objects simulated within the CGR environment are adjusted to behave according to at least one law of physics. For example, a CGR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. Depending on the situation (e.g., for accessibility reasons), the adjustment of the characteristics of the virtual object(s) in the CGR environment may be made in response to a representation of body movement (e.g., a voice command). People may use any one of these senses, including vision, hearing, touch, taste, and smell, to perceive and / or interact with CGR objects. For example, a person can perceive and / or interact with an audio object that creates an audio environment with a 3D or spatial spread that provides the perception of a point sound source in 3D space. In another example, audio objects can enable audio transparency that selectively incorporates ambient sound from the physical environment, with or without including computer-generated audio. In some CGR environments, people may only perceive and / or interact with audio objects.
[0046] Examples of CGR include virtual reality and mixed reality.
[0047] Virtual Reality: A virtual reality (VR) environment refers to an imitation environment designed to be fully based on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of the presence of the person within the computer-generated environment and / or through a simulation of a subset of the person's body movements within the computer-generated environment.
[0048] Mixed Reality: In contrast to a VR environment designed to be fully based on computer-generated sensory inputs, a mixed reality (MR) environment refers to an imitation environment designed to incorporate sensory inputs or their representations from the physical environment in addition to including computer-generated sensory inputs (e.g., virtual objects). On the virtual continuum, an MR environment is anywhere between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, the computer-generated sensory inputs can further respond to variations in the sensory inputs from the physical environment. Also, some electronic systems for presenting an MR environment may track the position and / or orientation with respect to the physical environment in order to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system can account for movement so that a virtual tree appears stationary relative to the physical ground.
[0049] Examples of mixed reality include augmented reality and augmented virtuality.
[0050] Augmented Reality: An augmented reality (AR) environment refers to an emulated environment in which one or more virtual objects are superimposed on a physical environment or its representation. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person can use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system synthesizes the image or video with virtual objects and presents the composite on the opaque display. A person uses this system to indirectly view the physical environment via the image or video of the physical environment and to perceive the virtual objects superimposed on the physical environment. As used herein, a video of a physical environment shown on an opaque display is referred to as a "pass-through video," meaning that the system uses one or more image sensors (singular or plural) to capture an image of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects into the physical environment or onto a physical surface, for example, as holograms, whereby a person can use the system to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulation environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, when providing a pass-through video, the system may transform one or more sensor images to map to a selected perspective (e.g., viewpoint) different from the perspective captured by the image sensor. As another example, a representation of the physical environment may be transformed by graphically modifying (e.g., magnifying) a portion thereof, thereby making the modified portion a modified version that represents the original captured image but is non-photorealistic. As a further example, a representation of the physical environment may be transformed by graphically removing or obscuring a portion thereof.
[0051] Extended Virtual: An extended virtual (AV) environment refers to an imitation environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but people with faces are realistically reproduced from images of physical people. As another example, virtual objects may adopt the shape or color of physical objects imaged by one or more imaging sensors. As a further example, virtual objects can adopt shadows that match the position of the sun in the physical environment.
[0052] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eye (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable controllers or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers (singular or plural) and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing an image or video of the physical environment and / or one or more microphones for capturing the audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards a person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scan light source, or any combination of these technologies. The medium may be an optical waveguide, hologram medium, optical coupler, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., physical settings / environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the scene 105. In some embodiments, the controller 110 is communicatively coupled to the display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH®, IEEE802.11x, IEEE802.16x, IEEE802.3x, etc.). In another example, the controller 110 is included within the housing (e.g., physical housing) of one or more of the display generation component 120 (e.g., an HMD, or a portable electronic device including a display and one or more processors, etc.), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0053] In some embodiments, the display generation component 120 is configured to provide a CGR experience (e.g., at least the visual component of the CGR experience) to the user. In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The display generation component 120 will be described in more detail below with reference to FIG. 3. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation component 120.
[0054] According to some embodiments, the display generation component 120 provides a CGR experience to the user while the user is virtually and / or physically present within the scene 105.
[0055] In some embodiments, the display generation component is worn on a part of the user's body (e.g., the head or hand). Thus, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds a device having a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally disposed within a housing worn on the user's head. In some embodiments, the handheld device is optionally disposed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content without the user wearing or holding the display generation component 120. Many of the user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a device on a handheld or tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface that indicates an interaction with CGR content triggered based on an interaction occurring within the space in front of a handheld or tripod-mounted device may be implemented in the same manner as an HMD where the interaction occurs within the space in front of the HMD and the response of the CGR content is displayed via the HMD. Similarly, a user interface that indicates an interaction with CRG content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)) may be implemented in the same manner as an HMD where the movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)) causes the interaction.
[0056] Although the relevant features of the operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features for the sake of brevity are not shown so as not to obscure more appropriate aspects of the exemplary embodiments disclosed herein.
[0057] FIG. 2 is a block diagram of an example of the controller 110 according to some embodiments. Although specific features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more appropriate aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE®, THUNDERBOLT®, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), BLUETOOTH®, ZIGBEE®, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0058] In some embodiments, one or more communication buses 204 include circuitry that interconnects system components and controls communication between the system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.
[0059] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile memory devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and a CGR experience module 240.
[0060] The operating system 230 includes instructions for processing various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). For that purpose, in various embodiments, the CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.
[0061] In some embodiments, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of FIG. 1 and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data acquisition unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0062] In some embodiments, the tracking unit 244 is configured to map the scene 105, at least for the display generation component 120 with respect to the scene 105 of FIG. 1, and optionally to track the position of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, the processing unit 244 includes a hand tracking unit 243 and / or an eye tracking unit 245. In some embodiments, the hand tracking unit 243 is configured to track the position of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand, with respect to the scene 105 of FIG. 1, with respect to the display generation component 120, and / or with respect to a coordinate system defined for the user's hand. The hand tracking unit 243 will be described in more detail below with reference to FIG. 4. In some embodiments, the eye tracking unit 245 is configured to track the position and movement of the user's line of sight (or, more generally, the user's eyes, face, or head) with respect to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hand)), or with respect to the CGR content displayed via the display generation component 120. The eye tracking unit 245 will be described in more detail below with reference to FIG. 5.
[0063] In some embodiments, the adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by the display generation component 120 and optionally by one or more of the output device 155 and / or the peripheral device 195. For that purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0064] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0065] Although the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 may be arranged within separate computing devices.
[0066] Furthermore, FIG. 2 is more intended to illustrate the functions of various features that may exist in a particular embodiment as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined, and some matters can be separated. For example, several functional modules separately shown in FIG. 2 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of particular functions and how functions are assigned among them, vary by embodiment and in some embodiments depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0067] Figure 3 is a block diagram of an example of a display generation component 120 according to some embodiments. While certain features are shown, those of ordinary skill in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more suitable aspects of the embodiments disclosed herein. For that purpose, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE®, THUNDERBOLT®, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH®, ZIGBEE®, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional inward and / or outward facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0068] In some embodiments, one or more communication buses 304 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0069] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission device display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays, such as diffractive, reflective, polarizing, holographic. For example, HMD 120 includes a single CGR display. In another example, HMD 120 includes a CGR display for each eye of the user. In some embodiments, one or more CGR displays 312 can present MR or VR content. In some embodiments, one or more CGR displays 312 can present MR or VR content.
[0070] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face, including the user's eyes (and may be referred to as an eye tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as a hand tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as a scene camera). The one or more optional image sensors 314 can include one or more RGB cameras, one or more infrared (IR) cameras, one or more event-based cameras, and / or the like (e.g., including a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor).
[0071] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or the non-transitory computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and a CGR presentation module 340.
[0072] The operating system 330 includes procedures for handling various basic system services and procedures for executing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to a user via one or more CGR displays 312. Thus, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.
[0073] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of FIG. 1. For that purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0074] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. For that purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0075] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a composite reality scene or a map of a physical environment in which computer-generated objects can be placed) based on media content data. For that purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0076] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For that purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0077] The data acquisition unit 342, the CGR presentation unit 344, and the CGR map generation unit 346, and the data transmission unit 348 are shown as being present on a single device (e.g., the display generation component 120 of FIG. 1), but in other embodiments, any combination of the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 may be arranged within separate computing devices.
[0078] Furthermore, FIG. 3 is more intended to illustrate the functions of various features that may exist in a particular embodiment as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined and some matters can be separated. For example, several functional modules separately shown in FIG. 3 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the particular division of specific functions and how functions are allocated therebetween, vary by embodiment and in some embodiments depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0079] FIG. 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (FIG. 1) is controlled by the hand tracking unit 243 (FIG. 2) to determine the position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 of FIG. 1 (e.g., relative to a part of the physical environment surrounding the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head)) and / or relative to a coordinate system defined with respect to the user's hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0080] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a resolution sufficient to distinguish the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either a zoom function or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with, or functions as, another image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.
[0081] In some embodiments, the image sensor 404 outputs a sequence of frames including 3D map data (and optionally color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and in response drives the display generation component 120. For example, the user can interact with software operating on the controller 110 by moving the hand 408 and changing the hand's posture.
[0082] In some embodiments, the image sensor 404 projects a spot pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spots of the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a predetermined reference plane at a particular distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the hand tracking device 440 can use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on a single or multiple cameras or other types of sensors.
[0083] In some embodiments, the hand tracking device 140 captures and processes a temporal sequence of depth maps that include the user's hand while the user is moving their hand (e.g., the entire hand or one or more fingers). Software operating on a processor within the image sensor 404 and / or the controller 110 processes the 3D map data to extract hand patch descriptors within these depth maps. The software compares these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the hand pose in each frame. The pose typically includes the 3D positions of the user's hand joints and fingertips.
[0084] Software can also analyze the trajectories of the hand and / or fingers over multiple frames in a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, such that patch-based pose estimation is only performed once every two (or more) frames, while tracking is used to detect changes in pose that occur over the remaining frames. Pose, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 in response to pose and / or gesture information, or perform other functions.
[0085] In some embodiments, the software may be downloaded in electronic form to the controller 110, for example, over a network, or alternatively, provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in a memory associated with the controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in FIG. 4 as a separate unit from the image sensor 440, by way of example, but some or all of the processing functions of the controller can be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the hand tracking device 402, or otherwise. In some embodiments, at least some of these processing functions are performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, a handheld device, or a head-mounted device), or using any other suitable computerized device such as a game console or a media player. The sensing function of the image sensor 404 can similarly be integrated with a computer or other computerized device controlled by the sensor output.
[0086] FIG. 4 further includes a schematic diagram of a depth map 410 captured by an image sensor 404 according to some embodiments. The depth map includes a matrix of pixels each having a respective depth value. The pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The luminance of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z - distance from the image sensor 404, and the tone becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment components (i.e., groups of adjacent pixels) of an image having characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and motion from frame to frame of a sequence of depth maps.
[0087] FIG. 4 also schematically shows a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on the background 416 of the hand segmented from the original depth map. In some embodiments, the hand (e.g., knuckles, fingertips, center of the palm, the end of the hand connected to the wrist, etc.), and optionally major feature points on the wrist or arm connected to the hand are identified and placed on the hand skeleton 414. In some embodiments, the positions and movements of these major feature points over a plurality of image frames are used by the controller 110 to determine, according to some embodiments, a hand gesture or the current state of the hand being performed by the hand.
[0088] FIG. 5 shows an exemplary embodiment of the eye tracking device 130 (FIG. 1). In some embodiments, the eye tracking device 130 is controlled by an eye tracking unit 245 (FIG. 2) to track the position and movement of the user's line of sight with respect to scene 105 or with respect to the CGR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed on a wearable frame, the head-mounted device includes both a component for generating CGR content for viewing by the user and a component for tracking the user's line of sight with respect to the CGR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, the eye tracking device 130 is optionally a device separate from the handheld device or the CGR chamber. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used with a display generation component worn on the head or a display generation component not worn on the head. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally a part of a non-head-mounted display generation component.
[0089] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that presents a frame including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and on which virtual objects can be displayed. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, whereby an individual can use the system to observe virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0090] As shown in FIG. 5, in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera), and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be directed at the user's eyes to directly receive reflected IR or NIR light from the light source, or alternatively, may be directed at a "hot" mirror disposed between the user's eyes and a display panel that reflects IR or NIR light from the eyes while allowing visual light to pass through to the eye tracking camera. The gaze tracking device 130 optionally captures an image of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the image to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a corresponding eye tracking camera and illumination source.
[0091] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility prior to delivery of the AR / VR device to the end user. The device-specific calibration process may be an automatic calibration process or a manual calibration process. The user-specific calibration process may include an estimation of a particular user's eye parameters, such as pupil position, central vision position, optical axis, visual axis, interpupillary distance, and the like. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 are determined, the images captured by the eye tracking camera are processed using a glint assist method to determine the user's current visual axis and viewpoint with respect to the display.
[0092] As shown in FIG. 5, the eye tracking device 130 (e.g., 130A or 130B) includes an eye tracking system including one or more eyepieces 520, at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) disposed on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 is positioned between the user's eye(s) 592 and a display 510 (e.g., a display panel on the left or right side of a head-mounted display, or a display of a handheld device, a projector, etc.), and may be directed toward a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (as shown, for example, at the top of FIG. 5), or may be directed toward the user's eye(s) 592 to receive the reflected IR or NIR light from the user's eye(s) 592 (as shown, for example, at the bottom of FIG. 5).
[0093] In some embodiments, the controller 110 renders an AR or VR frame 562 (e.g., the left and right frames of the left and right display panels) and provides the frame 562 to the display 510. The controller 110 uses the eye tracking input 542 from the eye tracking camera 540, for example, when processing the frame 562 for display, for various purposes. The controller 110 optionally estimates the user's viewing point on the display 510 based on the eye tracking input 542 obtained from the eye tracking camera 540 using a glint assist method or other suitable method. The viewing point estimated from the eye tracking input 542 is optionally used to determine the direction the user is currently looking.
[0094] Examples of several possible use cases of the user's current line of sight direction are described below, but this is not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined line of sight direction of the user. For example, the controller 110 may generate virtual content at a higher resolution in the central visual area determined from the user's current line of sight direction than in the peripheral area. As another example, the controller may position or move virtual content within the view at least partially based on the user's current line of sight direction. As another example, the controller may display specific virtual content within the view at least partially based on the user's current line of sight direction. As another exemplary use case in an AR application, the controller 110 can capture the physical environment of the CGR experience and direct an external camera to focus in the determined direction. The autofocus mechanism of the external camera can then focus on an object or surface within the environment that the user is currently viewing on the display 510. As another exemplary use case, the eyepiece 520 may be a focusable lens, and the line of sight tracking information is used by the controller to adjust the focus of the eyepiece 520 so that the virtual object the user is currently viewing has appropriate binocular coupling to match the convergence of the user's eyes 592. The controller 110 can utilize the line of sight tracking information to direct and adjust the focus of the eyepiece 520 so that the nearby object the user is viewing appears at the correct distance.
[0095] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye-tracking camera (e.g., eye-tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to a wearable housing. The light source emits light (e.g., IR or NIR light) towards the user's eye(s) 592. In some embodiments, the light source may be arranged in a ring or circularly around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 by way of example. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be employed.
[0096] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, so as not to introduce noise into the eye-tracking system. Note that the position and angle of the eye-tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 having a wide field of view (FOV) and a camera 540 having a narrow FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0097] An embodiment of a gaze tracking system as shown in FIG. 5 can be used, for example, in a computer-generated reality (including, for example, virtual reality and / or mixed reality) application to provide a user with an experience of computer-generated reality (including, for example, virtual reality, augmented reality, and / or augmented virtuality).
[0098] FIG. 6 shows a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., an eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the glint-assisted gaze tracking system uses prior information from the previous frame when analyzing the current frame to track the pupil contour and glint within the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint within the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.
[0099] As shown in FIG. 6, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0100] At 610, for the currently captured image, if the tracking state is yes, this method proceeds to element 640. At 610, if the tracking state is no, as shown at 620, the image is analyzed to detect the user's pupil and glint within the image. At 630, if the pupil and glint are detected successfully, the method proceeds to element 640. If not detected successfully, the method returns to element 610 to process the next image of the user's eye.
[0101] At 640, when proceeding from element 410, the current frame is analyzed to track the pupil and glint, based in part on the prior information from the previous frame. At 640, when proceeding from element 630, the tracking state is initialized based on the detected pupil and glint within the current frame. The result of the processing at element 640 is checked to confirm that the tracking or detection result is reliable. For example, the result can be checked to determine whether a sufficient number of glints for performing pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the result is not reliable, the tracking state is set to no and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's viewpoint.
[0102] FIG. 6 is intended to function as an example of an eye tracking technique that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye tracking techniques that currently exist or may be developed in the future can be used in the computer system 101, instead of, or in combination with, the glint-assisted gaze tracking technique described herein, to provide a user with a CGR experience according to various embodiments.
[0103] In the present disclosure, various input methods are described with respect to interaction with a computer system. It should be understood that when one example is provided using one input device or input method and another example is provided using another input device or input method, each example is compatible with the input device or input method described in the other example and can optionally be utilized. Similarly, various output methods are described with respect to interaction with a computer system. It should be understood that when one example is provided using one output device or output method and another example is provided using another output device or output method, each example is compatible with the output device or output method described in the other example and can optionally be utilized. Similarly, various methods are described with respect to interaction with a virtual environment or mixed reality environment via a computer system. It should be understood that when one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, each example is compatible with the method described in the other example and can optionally be utilized. Thus, the present disclosure discloses embodiments that are combinations of features of multiple examples without comprehensively listing all features of the embodiments in the description of each embodiment. User Interfaces and Related Processes
[0104] Here, attention is drawn to embodiments of a user interface (“UI”) and related processes that can be executed in a computer system such as a portable multifunctional device or a head-mounted device having a display generation component, one or more input devices, and (optionally) one or more cameras.
[0105] Figures 7A-7C show examples of input gestures (e.g., optionally, individual small motion gestures performed by moving a user's finger(s) relative to another finger(s) or a part(s) of the user's hand without requiring the user's entire hand or arm to move significantly away from their natural position(s) and orientation(s) immediately before or during the gesture to perform an action) for interacting with a virtual or mixed reality environment. The input gestures described with reference to FIGS. 7A-7C are used to illustrate the processes described below, including the process in FIG. 8.
[0106] In some embodiments, the input gestures described with reference to FIGS. 7A-7C are detected by analyzing data and signals captured by a sensor system (e.g., sensor 190 of FIG. 1, image sensor 314 of FIG. 3). In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras such as a motion RGB camera, an infrared camera, a depth camera). For example, the one or more imaging sensors are components of a computer system (e.g., computer system 101 of FIG. 1 (e.g., a portable electronic device 7100 or an HMD as shown in FIG. 7C)) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4 (e.g., a touch screen display that functions as a display and a touch sensing surface, a stereoscopic display, a display having a pass-through portion, etc.)), or provide data to the computer system. In some embodiments, the one or more imaging sensors include one or more rear cameras on a side of the device opposite the display of the device. In some embodiments, the input gestures are detected by a sensor system of a head-mounted system (e.g., a VR headset including a stereoscopic display that provides a left image of the user's left eye and a right image of the user's right eye). For example, one or more cameras that are components of the head-mounted system are attached to the front side and / or the lower side of the head-mounted system. In some embodiments, the one or more imaging sensors are arranged in the space in which the head-mounted system is used such that the imaging sensors capture images of the head-mounted system and / or the user of the head-mounted system (e.g., arranged around the head-mounted system at various positions in a room). In some embodiments, the input gestures are detected by a sensor system of a head-up device (e.g., a head-up display, an automotive windshield having the ability to display graphics, a window having the ability to display graphics, a lens having the ability to display graphics). For example, the one or more imaging sensors are attached to the inner surface of an automobile.In some embodiments, the sensor system includes one or more depth sensors (e.g., a sensor array). For example, the one or more depth sensors include one or more light-based (e.g., infrared) sensors and / or one or more acoustic-based (e.g., ultrasonic) sensors. In some embodiments, the sensor system includes one or more signal emitters such as a light emitter (e.g., an infrared emitter) and / or an acoustic emitter (e.g., an ultrasonic emitter). For example, while light (e.g., light from an infrared light emitter array having a predetermined pattern) is projected onto a hand (e.g., hand 7200 as described with respect to FIG. 71C), an image of the hand under the illumination of the light is captured by one or more cameras, and the captured image is analyzed to determine the position and / or configuration of the hand. In contrast to using signals from a touch sensing surface or other direct contact mechanism or proximity-based mechanism, by using signals from an image sensor directed at the hand to determine an input gesture, the user can freely choose to perform large movements or remain relatively stationary when providing an input gesture with the hand without experiencing the constraints imposed by a particular input device or input area.
[0107] Part (A) of FIG. 7A shows a tap input of the thumb 7106 on the index finger 7108 of the user's hand (e.g., on the side of the index finger 7108 adjacent to the thumb 7106). The thumb 7106 moves along the axis indicated by the arrow 7110, and this movement includes moving from the raised position 7102 to the touch-down position 7104 (e.g., the thumb 7106 touches the index finger 7110 and remains placed on the index finger 7110), and optionally, after the thumb 7106 touches the index finger 7110, moving from the touch-down position 7104 to the raised position 7102 within a threshold time (e.g., a tap time threshold). In some embodiments, the tap input is detected without the need to lift the thumb from the side of the index finger. In some embodiments, the tap input is detected according to the determination that the upward movement of the thumb follows the downward movement of the thumb and the thumb is in contact with the side of the index finger for less than the threshold time. In some embodiments, the tap-and-hold input is detected according to the determination that the thumb moves from the raised position 7102 to the touch-down position 7104 and stays at the touch-down position 7104 for at least a first threshold time (e.g., a tap time threshold or another time threshold longer than the tap time threshold). In some embodiments, the computer system requires the entire hand to remain substantially stationary at a position for at least the first threshold time in order to detect a tap-and-hold input by the thumb on the index finger. In some embodiments, the touch-and-hold input is detected without requiring the hand to remain substantially stationary (e.g., the entire hand can move while the thumb is placed on the side of the index finger). In some embodiments, the tap-and-hold drag input is detected when the thumb touches the side of the index finger and the entire hand moves while the thumb remains stationary on the side of the index finger.
[0108] Part (B) of FIG. 71A shows a push or flick input of the thumb 7116 crossing the index finger 7118 (e.g., from the palmar side to the dorsal side of the index finger). The thumb 7116 moves from the retracted position 7112 to the extended position 7114 along the axis indicated by the arrow 7120, crossing the index finger 7118 (e.g., crossing the middle phalanx of the index finger 7118). In some embodiments, the extension movement of the thumb involves an upward movement away from the side of the index finger, such as an upward flick input by the thumb. In some embodiments, the index finger moves in a direction opposite to the direction of the thumb while the thumb moves forward and upward. In some embodiments, the reverse flick input is performed by the thumb moving from the extended position 7114 to the retracted position 7112. In some embodiments, the index finger moves in a direction opposite to the direction of the thumb while the thumb moves backward and downward.
[0109] Part (C) of FIG. 7A shows a swipe input by the movement of the thumb 7126 along the index finger 7128 (e.g., along the side of the index finger 7128 adjacent to the thumb 7126 or along the palmar side). The thumb 7126 moves along the length of the index finger 7128 along the axis indicated by the arrow 7130 from the proximal position 7122 of the index finger 7128 (e.g., from the base phalanx of the index finger 7118 or in its vicinity) to the distal position 7124 (e.g., to the distal phalanx of the index finger 7118 or in its vicinity), and / or from the distal position 7124 to the proximal position 7122. In some embodiments, the index finger is optionally in an extended state (e.g., substantially straight) or a bent state. In some embodiments, the index finger moves between the extended state and the bent state while the thumb moves in the swipe input gesture.
[0110] Portion (D) of FIG. 7A shows tap inputs of thumb 7106 on various phalanges of various fingers (e.g., index finger, middle finger, ring finger, and optionally little finger). For example, thumb 7106 as shown in portion (A) moves from the raised position 7102 to a touch-down position as shown in any of 7130 - 7148 shown in portion (D). At touch-down position 7130, thumb 7106 is shown contacting a position 7150 on the proximal phalanx of index finger 7108. At touch-down position 7134, thumb 7106 is contacting a position 7152 on the middle phalanx of index finger 7108. At touch-down position 7136, thumb 7106 is contacting a position 7154 on the distal phalanx of index finger 7108.
[0111] At the touch-down positions shown at 7138, 7140, and 7142, thumb 7106 contacts positions 7156, 7158, and 7160 corresponding to the proximal phalanx, middle phalanx, and distal phalanx of the middle finger, respectively.
[0112] At the touch-down positions shown at 7144, 7146, and 7150, thumb 7106 contacts positions 7162, 7164, and 7166 corresponding to the proximal phalanx, middle phalanx, and distal phalanx of the ring finger, respectively.
[0113] In various embodiments, tap inputs by thumb 7106 on different portions of a different finger or different portions of two parallel fingers correspond to different inputs and trigger different actions in respective user interface contexts. Similarly, in some embodiments, different push or click inputs can be performed by the thumb across different fingers and / or different portions of a finger to trigger different actions in respective user interface contacts. Similarly, in some embodiments, different swipe inputs performed by the thumb along different fingers and / or in different directions (e.g., towards the distal or proximal end of a finger) trigger different actions in respective user interface contexts.
[0114] In some embodiments, the computer system processes tap input, flick input, and swipe input as different types of input based on the type of movement of the thumb. In some embodiments, the computer system processes input having different finger positions tapped, touched, or swiped by the thumb as different sub-input types (e.g., proximal, intermediate, distal sub-types, or index finger, middle finger, ring finger, or little finger sub-types) of a given input type (e.g., tap input type, flick input type, swipe input type, etc.). In some embodiments, the amount of movement performed by the moving finger (e.g., the thumb), and / or other measures of movement associated with the movement of the finger (e.g., speed, initial speed, end speed, duration, direction, movement pattern, etc.) are used to quantitatively affect the actions triggered by the finger input.
[0115] In some embodiments, the computer system recognizes combined input types that combine a series of movements by the thumb, such as tap-swipe input (e.g., the thumb swipes along the side of the finger after a touch-down to another finger), tap-flick input (e.g., the thumb flicks across the finger from the side of the palm to the back of the finger after a touch-down to another finger), double-tap input (e.g., two consecutive taps on the side of the finger at approximately the same position), etc.
[0116] In some embodiments, the gesture input is performed by the index finger instead of the thumb (e.g., the index finger performs a tap or a swipe on top of the thumb, or the thumb and the index finger move towards each other to perform a pinch gesture). In some embodiments, a movement of the wrist (e.g., a flick of the wrist in a horizontal or vertical direction) is performed immediately before, immediately after (e.g., within a threshold time), or simultaneously with the movement input of the finger to trigger an additional, different, or modified action in the current user interface context as compared to the movement input of the finger without the corrective input by the movement of the wrist. In some embodiments, a finger input gesture performed with the user's palm facing the user's face is treated as a different type of gesture than a finger input gesture performed with the user's palm facing away from the user's face. For example, a tap gesture performed with the user's palm facing the user is performed with an action (e.g., the same action) in response to a tap gesture performed with the user's palm facing away from the user's face, and an action with additional (or reduced) privacy protection is performed.
[0117] One type of finger input can be used to trigger the action type in the examples provided in this disclosure, but in other embodiments, other types of finger inputs are optionally used to trigger the same type of action.
[0118] FIG. 7B shows an exemplary user interface context showing a menu 7170 including user interface objects 7172-7194 in some embodiments.
[0119] In some embodiments, the menu 7170 is displayed in an augmented reality environment (e.g., floating in the air, or overlapping a physical object in a three-dimensional environment, corresponding to an action associated with the augmented reality environment or an action associated with the physical object). For example, the menu 7170 is displayed (e.g., overlaid) with at least a portion of a view of the physical environment captured by one or more rear cameras of the device 7100 on a display of the device (e.g., the device 7100 (FIG. 7C) or an HMD). In some embodiments, the menu 7170 is displayed on a transparent or translucent display of a device (e.g., a head-up display, or an HMD), through which the physical environment is visible. In some embodiments, the menu 7170 is displayed in a user interface that includes a pass-through portion surrounded by virtual content (e.g., a transparent or translucent portion through which the physical surroundings are visible, or a portion that displays a camera view of the surrounding physical environment). In some embodiments, the hand of a user performing a gesture input to execute an action in the augmented reality environment is visible to the user on the display of the device. In some embodiments, the hand of a user performing a gesture input to execute an action in the augmented reality environment is not visible to the user on the display of the device (e.g., the camera that provides the user with a view of the physical world has a different field of view than the camera that captures the user's finger input).
[0120] In some embodiments, the menu 7170 is displayed in the virtual reality environment (e.g., floating within the virtual space or overlapping a virtual surface). In some embodiments, the hand 7200 is visible in the virtual reality environment (e.g., an image of the hand 7200 captured by one or more cameras is rendered in the virtual reality setting). In some embodiments, a representation of the hand 7200 (e.g., a comic version of the hand 7200) is rendered in the virtual reality setting. In some embodiments, the hand 7200 is invisible in the virtual reality environment (e.g., omitted). In some embodiments, the device 7100 (FIG. 7C) is invisible in the virtual reality environment (e.g., when the device 7100 is an HMD). In some embodiments, an image of the device 7100 or a representation of the device 7100 is visible in the virtual reality environment.
[0121] In some embodiments, one or more of the user interface objects 7172 - 7194 are application launch icons (e.g., for performing an operation of launching a corresponding application). In some embodiments, one or more of the user interface objects 7172 - 7194 are controls for performing respective operations within an application (e.g., increasing volume, decreasing volume, playing, pausing, fast - forwarding, rewinding, starting communication with a remote device, ending communication with a remote device, relaying communication with a remote device, starting a game, etc.). In some embodiments, one or more of the user interface objects 7172 - 7194 are respective representations (e.g., avatars) of users of a remote device (e.g., for performing an operation of starting communication with each of the respective users of the remote device). In some embodiments, one or more of the user interface objects 7172 - 7194 are representations (e.g., thumbnails, two - dimensional images, or album covers) of media (e.g., images, virtual objects, audio files, and / or video files). For example, by activating a user interface object that is a representation of an image, the image is displayed at a position corresponding to a surface detected by one or more cameras and displayed in a computer - generated reality view (e.g., a position corresponding to a surface within a physical environment or a position corresponding to a surface displayed in a virtual space).
[0122] When the thumb of the hand 7200 performs the input gesture described with respect to FIG. 7A, the operation corresponding to the menu 7170 is executed according to the position and / or type of the detected input gesture. For example, in response to an input including a movement of the thumb along the y-axis (e.g., a movement from a proximal position on the index finger to a distal position on the index finger as described with respect to part (C) of FIG. 7A), the current selection indicator 7198 (e.g., a selector object or a movable visual effect such as highlighting an object by a contour or a change in the appearance of the object) is repeatedly moved to the right from the item 7190 to the subsequent user interface object 7192. In some embodiments, in response to an input including a movement of the thumb along the y-axis from a distal position on the index finger to a proximal position on the index finger, the current selection indicator 7198 is repeatedly moved to the left from the item 7190 to the previous user interface object 7188. In some embodiments, in response to an input including a tap input of the thumb on the index finger (e.g., a movement of the thumb along the z-axis as described with respect to part (A) of FIG. 7A), the currently selected user interface object 7190 is activated, and the operation corresponding to the currently selected user interface object 7190 is executed. For example, the user interface object 7190 is an application launch icon, and while the user interface object 7190 is selected, in response to the tap input, the application corresponding to the user interface object 7190 is launched and displayed on the display. In some embodiments, in response to an input including a movement of the thumb along the x-axis from a retracted position to an extended position, the current selection indicator 7198 moves upward from the item 7190 to the upper user interface object 7182. Other types of finger inputs can be provided for one or more of the user interface objects 7172-7194, and optionally, other types of operations corresponding to the user interface object(s) are executed according to the input.
[0123] FIG. 7C shows a visual indication of menu 7170 that is visible in a composite reality view (e.g., an augmented reality view of a physical environment) displayed by a computer system (e.g., device 7100 or HMD). In some embodiments, hand 7200 within the physical environment is visible within the displayed augmented reality view (e.g., as part of a view of the physical environment captured by a camera), as illustrated at 7200'. In some embodiments, hand 7200 is visible through a transparent or translucent display surface on which menu 7170 is displayed (e.g., device 7100 is a head-up display or HMD with a pass-through portion).
[0124] In some embodiments, as shown in FIG. 7C, menu 7170 is displayed at a position within the composite reality environment that corresponds to a predetermined portion of the user's hand (e.g., the tip of the thumb) and has an orientation corresponding to the orientation of the user's hand. In some embodiments, as the user's hand moves (e.g., moves laterally or rotates) relative to the physical environment (e.g., the user's hand, or the user's eyes, or a camera that captures physical objects or walls surrounding the user), menu 7170 is shown to move within the composite reality environment with the user's hand. In some embodiments, menu 7170 moves in accordance with the movement of the user's line of sight directed towards the composite reality environment. In some embodiments, menu 7170 is displayed at a fixed position on the display regardless of the view of the physical environment shown on the display.
[0125] In some embodiments, menu 7170 is displayed on the display in response to detecting a ready position of the user's hand (e.g., the thumb placed on the side of the index finger). In some embodiments, user interface objects displayed in response to detecting a ready-position hand vary according to the current user interface context and / or the position of the user's line of sight within the composite reality environment.
[0126] Figures 7D - 7E show hand 7200 in an exemplary not - ready state configuration (e.g., a rest configuration) (Figure 7D) and an exemplary ready state configuration (Figure 7E) according to some embodiments. The input gestures described with reference to Figures 7D - 7E are used to illustrate the processes described below, including the process in Figure 9.
[0127] In FIG. 7D, the hand 7200 is shown in an exemplary non-ready-complete state configuration (e.g., a rest configuration (e.g., the hand is relaxed or in any state, and the thumb 7202 is not placed on the index finger 7204)). In an exemplary user interface context, a container object 7206 (e.g., an application dock, folder, control panel, menu, platter, preview, etc.) that includes user interface objects 7208, 7210, and 7212 (e.g., application icons, media objects, controls, menu items, etc.) is displayed in a three-dimensional environment (e.g., a virtual environment or a mixed reality environment) by a display generation component (e.g., the display generation component 120 in FIGS. 1, 2, and 4 (e.g., a touch screen display, a stereoscopic projector, a head-up display, an HMD, etc.)) of a computer system (e.g., the computer system 101 in FIG. 1 (e.g., a system including the device 7100, an HMD, or other devices)). In some embodiments, the rest configuration of the hand is an example of a hand configuration that is not a ready-complete state configuration. For example, other hand configurations that are not necessarily relaxed and stationary and do not meet the criteria for detecting a ready-complete state gesture (e.g., the thumb is placed on the index finger (e.g., the middle phalanx of the index finger)) are also classified as non-ready-complete state configurations. For example, when the user is waving their hand in the air, or holding an object, or clenching a fist, etc., in some embodiments, the computer system determines that the user's hand does not meet the criteria for detecting a ready-complete state of the hand and determines that the user's hand is in a non-ready-complete state configuration. In some embodiments, the criteria for detecting a ready-complete state configuration of the hand include detecting that the user has changed their hand configuration, and as a result of that change, the user's thumb is placed on a predetermined portion (e.g., the middle phalanx of the index finger) of the user's index finger. In some embodiments, the criteria for detecting a ready-complete state configuration require that the user's thumb be placed on a predetermined portion of the user's index finger for at least a first threshold time as a result of a change in the hand gesture for the computer system to recognize that the hand is in a ready-complete state configuration.In some embodiments, after the user enters the ready state configuration without changing their hand configuration, if no valid input gesture is provided for at least a second threshold time, the computer system processes the current hand configuration as a non-ready state configuration. In the computer system, the user needs to change their current hand configuration and then return to the ready state configuration to re-recognize it. In some embodiments, the ready state configuration is user-configurable and user-customizable, for example, by the user showing the intended ready state configuration of the hand and optionally the acceptable variation range of the ready state configuration to the computer system within the gesture setting environment provided by the computer system.
[0128] In some embodiments, the container 7206 is displayed in the mixed reality environment (as shown, for example, in FIGS. 7D and 7E). For example, the container 7206 is displayed on the display of the device 7100 that has at least a portion of the view of the physical environment captured by one or more rear cameras of the device 7100. In some embodiments, the hand 7200 within the physical environment is also visible in the displayed mixed reality environment, for example, having an actual spatial relationship between the hand and the physical environment represented by the displayed view of the mixed reality environment, as shown at 7200b. In some embodiments, the container 7206 and the hand 7200 are displayed remotely from the user and in relation to the physical environment displayed via a live feed of a camera coupled to a remote physical environment. In some embodiments, the container 7206 is displayed on a transparent or translucent display of a device where the physical environment surrounding the user (including the hand 7200 as shown at 7200b) is visible.
[0129] In some embodiments, the container 7206 is displayed in a virtual reality environment (e.g., floating within a virtual space). In some embodiments, the hand 7200 is visible in a virtual reality setting (e.g., an image of the hand 7200 captured by one or more cameras is rendered in the virtual reality environment). In some embodiments, the representation of the hand 7200 is visible in the virtual reality environment. In some embodiments, the hand 7200 is invisible in the virtual reality environment (e.g., omitted). In some embodiments, the device 7100 is invisible in the virtual reality environment. In some embodiments, an image of the device 7100 or a representation of the device 7100 is visible in the virtual reality environment.
[0130] In some embodiments, while the hand 7200 is not in a ready state configuration (e.g., in any non-ready state configuration or has stopped remaining in the ready state configuration, such as because the hand gesture changes or cannot provide a valid input gesture within a threshold time to enter the ready state configuration), the computer system does not perform input gesture recognition for executing operations in the current user interface context (other than recognizing whether the hand has entered the ready state configuration). As a result, in response to an input gesture performed by the hand 7200 (e.g., a tap of the thumb on the index finger, including movement of the thumb along the axis indicated by the arrow 7110 as described with respect to part (A) of FIG. 7A, a movement of the thumb across the index finger along the axis indicated by the arrow 7120 as described with respect to part (B) of FIG. 7A, and / or a movement of the thumb on the index finger along the axis indicated by the arrow 7130 as described with respect to part (C) of FIG. 7A), no operation is performed. In other words, the computer system sets the user to the ready state configuration (e.g., changes from a non-ready state configuration to a ready state configuration) in order to recognize the input gesture as valid and perform corresponding operations in the current user interface context, and then needs to provide a valid input gesture for the current user interface context (e.g., within a threshold time when the hand enters the ready state configuration). In some embodiments, if a valid input gesture is detected without first detecting a hand in the ready state configuration, the computer system performs a certain type of operation (e.g., interacts with (e.g., scrolls or activates) the currently displayed user interface object) and prohibits other types of operations (e.g., calling a new user interface, triggering system-level operations (e.g., navigating to a multitasking user interface or an application launch user interface, activating a device function control panel, etc.)).These protections serve to prevent and reduce inadvertent and unintentional action triggers and avoid unnecessary restrictions on the user's free hand movement if the user does not wish to operate or perform a particular type of operation within the current user interface context. Additionally, imposing small individual movement requirements on the ready state configuration tends to reduce the user's sense of discomfort when interacting with the user interface in a social environment, without imposing an excessive physical burden on the user (e.g., by causing the user to move their arm or hand excessively).
[0131] The user interface objects 7208 - 7212 of the container 7206 include, for example, one or more application launch icons, one or more control parts for executing operations within an application, one or more representations of users of a remote device, and / or one or more representations of media (e.g., as described above with respect to the user interface objects 7172 - 7194). In some embodiments, when a user interface object is selected and an input gesture is detected without the hand being first discovered in a ready - to - go configuration by the computer system, the computer system performs a first operation according to the input gesture with respect to the selected user interface object (e.g., launches the application corresponding to the selected application icon, changes the control value of the selected control part, initiates communication with the user regarding the selected user representation, starts playing the media item corresponding to the representation of the selected media item). When the user interface object is selected and the same input gesture is detected with the hand first discovered in a ready - to - go configuration by the computer system, the computer system performs a second operation different from the first operation (e.g., the second operation is a system operation not specific to the currently selected user interface object (e.g., the system operation includes displaying system affordances in response to the hand being discovered in a ready - to - go configuration and launching a system menu in response to the input gesture)). In some embodiments, by placing the hand in a ready - to - go configuration, a specific input gesture (e.g., a thumb flick gesture) not paired with any function within the current user interface context is enabled, and when the newly enabled input gesture is detected after the hand is discovered in a ready - to - go configuration, the computer system is made to execute additional functions associated with the newly enabled input gesture.In some embodiments, the computer system optionally displays a user interface indication (e.g., additional options, system affordances, or a system menu) in response to detecting a ready state configuration of the hand, and enables the user to interact with the user interface indication or trigger additional functionality using the newly enabled input gesture (e.g., a thumb flick gesture detected when a system affordance is displayed causes the system menu to be displayed, and a thumb flick gesture detected when the system menu is displayed causes navigation by the system menu or an expansion of the system menu).
[0132] In FIG. 7E, a hand 7200 is shown in a ready state configuration (e.g., a thumb 7202 is placed on an index finger 7204). In accordance with the determination that the hand 7200 has moved to the ready state configuration, the computer system displays a system affordance icon 7214 (e.g., within a region of the mixed reality environment corresponding to the tip of the thumb). The system affordance icon 7214 indicates a region that can display and / or access one or more user interface objects (e.g., a menu of application icons corresponding to different applications, a menu of currently open applications, a device control user interface, etc.) in response to an input gesture performed by the hand 7200 (e.g., as described below with respect to FIG. 7F). In some embodiments, while the system affordance icon 7214 is displayed, the computer system performs an operation in response to an input gesture performed by the hand 7200 (e.g., as described below with respect to FIG. 7F).
[0133] In some embodiments, the movement of the hand 7200 from the not-ready state configuration to the ready state configuration is detected by analyzing data captured by a sensor system (e.g., an image sensor, or other sensors such as a motion sensor, a touch sensor, a vibration sensor, etc.) as described above with respect to FIGS. 7A-7C. In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras associated with a portable device or a head-mounted system), one or more depth sensors, and / or one or more light emitters.
[0134] In some embodiments, the system affordance icon 7214 is displayed in the mixed reality environment. For example, the system affordance icon 7214 is displayed by a display generation component (e.g., the display of the device 7100 or the HMD) together with at least a portion of a view of the physical environment captured by one or more cameras of the computer system (e.g., one or more rear cameras of the device 7100, or the front or bottom cameras of the HMD). In some embodiments, the system affordance icon 7214 is displayed on a transparent or semi-transparent display of a device (e.g., a head-up display, or an HMD with a passthrough portion) through which the physical environment is visible. In some embodiments, the system affordance icon 7214 is displayed in a virtual reality environment (e.g., floating within a virtual space).
[0135] In some embodiments, the computer system stops displaying the system affordance icon 7214 in response to detecting that the hand gesture has changed and is no longer in a ready state configuration without providing a valid input gesture for the current user interface context. In some embodiments, the computer system stops displaying the system affordance icon 7214 and determines that the criteria for detecting a ready state configuration are no longer met in response to detecting that the hand has remained in a ready state gesture without providing a valid input gesture for a threshold time. In some embodiments, after stopping the display of the system affordance icon 7214, in accordance with a determination that the criteria for detecting a ready state configuration are again met due to a change in the user's hand gesture, the computer system redisplay the system affordance icon (e.g., at the tip of the thumb at the new hand position).
[0136] In some embodiments, two or more ready state configurations of the hand are optionally defined and recognized by the computer system, and each ready state configuration of the hand causes the computer system to display a different type of affordance and enables a different set of input gestures and / or actions to be performed in the current user interface context. For example, a second ready state configuration optionally causes all fingers to cooperate in a fist-like manner such that the thumb is placed on a finger other than the index finger. When the computer system detects this second ready state configuration, the computer system displays a different system affordance icon from icon 7214, and a subsequent input gesture (e.g., a thumb swipe across the index finger) causes the computer system to initiate a system shutdown action or display a menu of power options (e.g., shutdown, sleep, suspend, etc.).
[0137] In some embodiments, the system affordance icon displayed in response to the computer detecting a hand in a ready state configuration is a home affordance that shows a selection user interface including a plurality of currently installed applications, and is displayed in response to the detection of a predetermined input gesture (e.g., a thumb flick input, a thumb push input, or other input as described with respect to FIGS. 7A-7C). In some embodiments, an application dock including a plurality of application icons for launching each application is displayed in response to the computer detecting a hand in a ready state configuration, and a subsequent activation or hand selection input gesture causes the application corresponding to the computer system to be launched.
[0138] FIGS. 7F-7G provide various examples of operations performed in response to input gestures detected in a state where a hand is discovered in a ready state configuration according to various embodiments. FIGS. 7F-7G illustrate different operations performed in response to the user's line of sight, but it should be understood that in some embodiments, the line of sight is not an element necessary to trigger the execution of their various functions in the detection of ready state gestures and / or input gestures. In some embodiments, the combination of the user's line of sight and hand configuration is used by the computer system in conjunction with the current user interface context to determine which operation is performed.
[0139] FIG. 7F shows an exemplary input gesture performed by hand in a ready state configuration according to some embodiments, and an exemplary response of the displayed three-dimensional environment (e.g., virtual reality environment or mixed reality environment). In some embodiments, the computer system (e.g., device 7100 or HMD) displays a system affordance (e.g., system affordance icon 7214) at the tip of the thumb to indicate that the hand is in a ready state configuration and that system gestures (e.g., a thumb flick gesture to display a dock or system menu, a thumb tap gesture to activate a voice-based assistant) can trigger a predetermined system action in addition to other input gestures (e.g., a thumb swipe input to scroll a user interface object) already available in the current user interface context. In some embodiments, the system affordance moves in accordance with the movement of the entire hand and / or the movement of the thumb in space when providing gesture input, such that the position of the system affordance remains fixed relative to a predetermined portion of the hand (e.g., the tip of the thumb of hand 7200).
[0140] According to some embodiments, portions (A)-(C) of FIG. 7F show a three-dimensional environment 7300 (e.g., virtual reality environment or mixed reality environment) displayed by a display generation component of a computer system (e.g., the touch screen display of device 7100 or a stereoscopic projector or the display of an HMD). In some embodiments, device 7100 is a handheld device (e.g., a mobile phone, tablet, or other mobile electronic device) that includes a display, a touch-sensitive display, etc. In some embodiments, device 7100 represents a wearable headset that includes a head-up display, a head-mounted display, etc.
[0141] In some embodiments, the three-dimensional environment 7300 is a virtual reality environment that includes virtual objects (e.g., user interface objects 7208, 7210, and 7212). In some embodiments, the virtual reality environment does not correspond to the physical environment in which the device 7100 is located. In some embodiments, the virtual reality environment corresponds to the physical environment (e.g., at least some of the virtual objects are displayed at positions in the virtual reality environment that correspond to the positions of physical objects in the physical environment determined using one or more cameras of the device 7100). In some embodiments, the three-dimensional environment 7300 is a mixed reality environment. In some embodiments, the device 7100 includes one or more cameras configured to continuously provide a live view of at least a portion of the surrounding physical environment within the field of view of the one or more cameras of the device 7100, and the mixed reality environment corresponds to the portion of the surrounding physical environment within the field of view of the one or more cameras of the device 7100. In some embodiments, the mixed reality environment at least partially includes the live view of the one or more cameras of the device 7100. In some embodiments, the mixed reality environment is displayed (e.g., including one or more virtual objects that are superimposed, overlapped, or replaced) instead of the live camera view at a position in the three-dimensional environment 7300 that corresponds to the position of a physical object in the physical environment (e.g., based on the position of the physical object in the physical environment determined using one or more cameras or using the live view of the one or more cameras of the device 7100). In some embodiments, the display of the device 7100 includes a head-up display that is at least partially transparent (e.g., having a transparency threshold less than 25%, 20%, 15%, 10%, or less than 5%, or having a pass-through portion) such that the user can view at least a portion of the surrounding physical environment through at least a partially transparent region of the display.In some embodiments, the three-dimensional environment 7300 includes one or more virtual objects (e.g., containers 7206 including user interface objects 7208, 7210, and 7212) displayed on a display. In some embodiments, the three-dimensional environment 7300 includes one or more virtual objects displayed on a transparent region of the display such that the three-dimensional environment appears superimposed on a portion of the surrounding physical environment visible through the transparent region of the display. In some embodiments, each of the one or more virtual objects is displayed at a location within the three-dimensional environment 7300 corresponding to the location of a physical object within the physical environment (e.g., as determined using one or more cameras of a device 7100 that monitors a portion of the physical environment visible through the transparent region of the display) such that each virtual object is displayed in place of the respective physical object (e.g., obscuring and replacing the view of the respective physical object).
[0142] In some embodiments, a sensor system of a computer system (e.g., one or more cameras of device 7100 or an HMD) tracks the position and / or movement of one or more features of a user, such as the user's hand. In some embodiments, the position and / or movement of the user's hand(s) (e.g., finger(s)) functions as an input to the computer system (e.g., device 7100 or an HMD). In some embodiments, the user's hand is within the field of view of one or more cameras of the computer system (e.g., device 7100 or an HMD), and the position and / or movement of the user's hand is tracked by the sensor system of the computer system (e.g., the control unit of device 7100 or an HMD) as an input to the computer system (e.g., device 7100 or an HMD), but the user's hand is not shown in the three-dimensional environment 7300 (e.g., the three-dimensional environment 7300 does not include a live view from one or more cameras, the hand is edited outside of the live view of one or more cameras, or the user's hand is within the field of view of one or more cameras outside of the portion of the field of view that is displayed in the live view of the three-dimensional environment 7300). In some embodiments, similar to the example shown in FIG. 7F, a hand 7200 (e.g., a representation of the user's hand or a representation of a portion of the hand within the field of view of one or more cameras of device 7100) is visible in the three-dimensional environment 7100 (e.g., displayed as a rendered representation, displayed as part of a live camera view, or visible through a pass-through portion of the display). In FIG. 7F, the hand 7200 is detected to be in a ready state configuration (e.g., thumb at the center of the index finger) for providing a gesture input (e.g., a hand gesture input). In some embodiments, the computer system (e.g., device 7100 or an HMD) determines that the hand 7200 is in the ready state configuration by performing image analysis on the live view of one or more cameras. Further details regarding the ready state configuration of the hand and input gestures are provided in at least FIGS. 7A-7E and the accompanying description and are not repeated here for the sake of brevity.
[0143] According to some embodiments, part (A) of FIG. 7F shows a first type of input gesture (e.g., a thumb flick gesture or a thumb push gesture) that includes the movement of the thumb of hand 7200 across a portion of the index finger of hand 7200 along an axis indicated by arrow 7120 (e.g., across the middle phalanx from the palmar side to the dorsal side of the index finger, as shown by the transition of the hand configuration from A(1) to A(2) in FIG. 7F). As described above, in some embodiments, the movement of hand 7200 while performing a gesture as shown in FIG. 7F (e.g., including the movement of the entire hand and the relative movement of individual fingers during the gesture), and the position of hand 7200 during the gesture (e.g., including the position of the entire hand and the relative positions of individual fingers) are tracked by one or more cameras of device 7100. In some embodiments, device 7100 detects a gesture performed by hand 7200 by performing image analysis on the live view of one or more cameras. In some embodiments, in accordance with a determination that a thumb flick gesture starting from a hand ready state configuration is being provided by the hand, the computer system performs a first action corresponding to the thumb flick gesture (e.g., a system action such as displaying menu 7170 that includes multiple application launch icons, or an action corresponding to the current user interface context that would not be enabled if the hand was not first discovered in the ready state configuration). In some embodiments, additional gestures are provided by hand 7200 to interact with the menu. For example, a subsequent thumb flick gesture after the display of menu 7170 drives menu 7170 into a three-dimensional environment and displays it in an enhanced state (e.g., using a more animated representation of menu items) in a virtual space. In some embodiments, a subsequent thumb swipe gesture horizontally scrolls a selection indicator within the currently selected row of the menu, and a subsequent thumb push or thumb pull gesture scrolls the selection indicator up and down across different rows of the menu. In some embodiments, a subsequent thumb tap gesture deactivates the currently selected menu item (e.g., an application icon) launched in a three-dimensional environment.
[0144] According to some embodiments, part (B) of FIG. 7F shows a second type of gesture (e.g., a thumb swipe gesture) that includes movement of the thumb of hand 7200 along the length of the index finger of hand 7200 along an axis indicated by arrow 7130 (as shown, for example, by the transition of the hand configuration to B(1)-B(2) in FIG. 7F). In some embodiments, the gesture enables interaction with virtual objects (e.g., virtual objects 7208, 7210, and 7212 within container object 7206) in the current user interface context, regardless of whether the hand was initially found in the ready state configuration. In some embodiments, different types of interactions with container object 7206 are enabled depending on whether the thumb swipe gesture was initiated from the ready state configuration. In some embodiments, in accordance with the determination that the thumb swipe gesture was provided by a hand starting from the ready state configuration, the computer system performs a second action corresponding to the thumb swipe gesture (e.g., scrolls the view of container object 7206 to display one or more virtual objects within container object 7206 that were initially invisible to the user, or another action corresponding to the current user interface context that is not enabled if the hand was not initially found in the ready state configuration). In some embodiments, additional gestures are provided by hand 7200 to interact with the container object or perform system operations. For example, a subsequent thumb flick gesture causes menu 7170 to be displayed (as shown, for example, in part (A) of FIG. 7F). In some embodiments, subsequent thumb swipe gestures in different directions scroll the view of container object 7206 in the opposite direction. In some embodiments, a subsequent thumb tap gesture activates the currently selected virtual object within container 7206. In some embodiments, if the thumb swipe gesture was not initiated from the ready state configuration, in response to the thumb swipe gesture, a selection indicator is shifted through the virtual objects within container 7206 in the direction of movement of the thumb of hand 7200.
[0145] According to some embodiments, part (C) of FIG. 7F represents a third type of gesture input (e.g., a thumb tap gesture) (e.g., as indicated by the transition of the hand configuration from C(1) to C(2) in FIG. 7F) (e.g., by moving the thumb downward from an elevated position along the axis indicated by arrow 7110) (e.g., a tap of the thumb of hand 7200 on a predetermined portion (e.g., the middle phalanx) of the index finger of hand 7200). In some embodiments, to complete a thumb tap gesture, it is necessary to separate the thumb from the index finger. In some embodiments, the gesture enables interaction with virtual objects (e.g., virtual objects 7208, 7210, and 7212 within container object 7206) in the current user interface context, regardless of whether the hand was initially discovered in a ready state configuration. In some embodiments, different types of interactions with container object 7206 are enabled depending on whether the thumb tap gesture was initiated from a ready state configuration of the hand. In some embodiments, in accordance with a determination that a thumb tap gesture initiated from a ready state configuration of the hand was provided by the hand (e.g., the gesture includes moving the thumb away from the index finger and upward prior to tapping the thumb on the index finger), the computer system performs a third action corresponding to the thumb tap gesture (e.g., activation of a voice-based assistant 7302 or a communication channel (e.g., a voice communication application)), or another action corresponding to the current user interface context that is not enabled if the hand was not initially discovered in a ready state configuration. In some embodiments, additional gestures are provided by hand 7200 to interact with a voice-based assistant or a communication channel. For example, a subsequent thumb flick gesture drives a voice-based assistant or a voice communication user interface from an adjacent position of the user's hand to a more distant position within the space in a three-dimensional environment. In some embodiments, a subsequent thumb swipe gesture scrolls through different pre-set functions of a voice-based assistant or scrolls through a list of potential recipients of a voice communication channel.In some embodiments, a subsequent thumb tap gesture causes a voice-based assistant or a voice communication channel to be dismissed. In some embodiments, if the thumb tap gesture does not start from the ready state configuration, in response to the thumb tap gesture in part (C) of FIG. 7E, an operation available in the current user interface context is activated (e.g., the currently selected virtual object is activated).
[0146] The example shown in FIG. 7F is merely illustrative. By providing additional and / or different operations in response to the detection of an input gesture starting from the ready state configuration of the hand, the user can perform additional functions without cluttering the user interface at the control unit, and by reducing the number of user inputs for performing those functions, the user interface can be made more efficient and the time taken during the interaction between the user and the device can be saved.
[0147] The user interface interactions shown in FIGS. 7A-7F are described regardless of the position of the user's line of sight. In some embodiments, the interaction does not depend on the position of the user's line of sight or the exact position of the user's line of sight within a portion of the three-dimensional environment. However, in some embodiments, the line of sight is utilized to modify the response behavior of the system, and the actions and user interface feedback are changed according to the different positions of the user when user input is detected. FIG. 7G shows an exemplary gesture performed by a hand in a ready state and an exemplary response of the three-dimensional environment displayed according to the user's line of sight, according to some embodiments. For example, the left column of the figure (e.g., portions A-0, A-1, A-2, and A-3) shows that one or more gesture inputs (e.g., a thumb flick gesture, a thumb swipe gesture, a thumb tap gesture, or a sequence of two or more of the above) are in a ready state configuration (e.g., the thumb is placed on the index finger) (e.g., the ready state configuration described in connection with FIGS. 7E and 7F) while the user's line of sight is focused on the user's hand (e.g., as shown in FIG. 7G, portion A-0), showing an exemplary scenario provided by the hand starting from. The right column of the figure (e.g., portions B-0, B-1, B-2, and B-3) shows that one or more gesture inputs (e.g., a thumb flick gesture, a thumb swipe gesture, a thumb tap gesture, or a sequence of two or more of them) are in a ready state configuration (e.g., the thumb is placed on the index finger) (e.g., the ready state configuration described in connection with FIGS. 7E and 7F), optionally, while the user's line of sight is focused on a user interface environment other than the user's hand in the ready state configuration (e.g., a user interface object or a physical object within the three-dimensional environment) (e.g., as shown in FIG. 7G, portion A-0), showing an exemplary scenario provided by the hand starting from.In some embodiments, the user interface responses and interactions described with respect to FIG. 7G are performed with respect to other physical objects or predetermined portions thereof (e.g., as contrasted with the user's hand providing the input gesture, the upper or front surface of the housing of a physical media player device, a physical window in the wall of a room, a physical controller device, etc.), such that while the user's line of sight is focused on the other physical object or predetermined portion thereof, special interactions are enabled for input gestures provided by the hand that commence from the ready state configuration. The figures in the left and right columns of the same row of FIG. 7G (e.g., A-1 and B-1, A-2 and B-2, A-3, and B-3) show different user interface responses to the same input gesture provided in the same user interface context, depending on whether the user's line of sight is focused on the user's hand (or other physical object defined as a physical object being controlled or controlled by the computer system). The input gestures described with reference to FIG. 7G are used to exemplify the processes described below, including the process in FIG. 10.
[0148] In some embodiments, as shown in FIG. 7G, the system affordance (e.g., affordance 7214) is optionally displayed at a predetermined three-dimensional position corresponding to the hand in the ready state configuration. In some embodiments, the system affordance is always displayed at a predetermined position (e.g., a static position or a dynamically determined position) whenever the hand is determined to be in the steady state configuration. In some embodiments, the dynamically determined position of the system affordance is fixed relative to the user's hand in the ready state configuration, for example, when the entire user's hand moves while remaining in the ready state configuration. In some embodiments, the system affordance is displayed (e.g., at a static position at or near the tip of the thumb) in response to the user's line of sight being directed towards a predetermined physical object (e.g., the user's hand in the ready state configuration, or another predetermined physical object in the environment) while the user's hand is maintained in the ready state configuration, and the display is stopped in response to the user's hand moving away from the ready state configuration and / or the user's line of sight moving away from the predetermined physical object. In some embodiments, the computer system displays a system affordance having a first appearance (e.g., an enlarged and prominent appearance) in response to the user's line of sight being directed towards a predetermined physical object (e.g., the user's hand in the ready state configuration, or another predetermined physical object in the environment) while the user's hand is maintained in the ready state configuration, and displays a system affordance having a second appearance (e.g., a reduced and unobtrusive appearance) in response to the user's hand moving away from the ready state configuration and / or the user's line of sight moving away from the predetermined physical object. In some embodiments, the system affordance is displayed (e.g., at a static position at or near the tip of the thumb) in response to an indication that the user is ready to provide an input (e.g., the user's hand rising from a previous elevation relative to the body while the user's hand is maintained in the ready state configuration), regardless of whether the user's line of sight is directed towards the user's hand, and the display is stopped in response to the user's hand descending from the raised state.In some embodiments, the system affordance changes the appearance of the system affordance (e.g., from a simple indicator to an object menu) in response to an indication that the user is ready to provide input (e.g., while the user's hand is in a ready state configuration and / or the user's line of sight is focused on the user's hand or a predetermined physical object, and the user's hand is rising above the forward altitude relative to the body), and returns the appearance of the system affordance to its original state (e.g., to a simple indicator) in response to the cessation of an indication that the user is ready to provide input (e.g., the user's hand is descending from the raised state and / or the user's line of sight is moving away from the user's hand or a predetermined physical object). In some embodiments, an indication that the user is ready to provide input includes one or more of the user's fingers touching a physical controller or the user's hand (e.g., the index finger is on the controller or the thumb is on the index finger), the user's hand rising from a lower position to a higher position relative to the user's body (e.g., there is a rotation of the wrist upward of the hand in the ready state configuration or a flexion of the elbow of the hand in the ready state configuration), changing the hand configuration to the ready state configuration, and the like. In some embodiments, the physical position of the predetermined physical object compared to the position of the user's line of sight in these embodiments is static with respect to the three-dimensional environment (e.g., also referred to as "fixed to the world"). For example, the system affordance is displayed on a wall within the three-dimensional environment. In some embodiments, the physical position of the predetermined physical object compared to the position of the user's line of sight in these embodiments is static with respect to the display (e.g., a display generation component) (e.g., also referred to as "fixed to the user"). For example, the system affordance is displayed at the bottom of the display or the user's field of view. In some embodiments, the physical position of the predetermined physical object compared to the position of the user's line of sight in these embodiments is static with respect to a moving part of the user (e.g., the user's hand) or a moving part of the physical environment (e.g., a car moving on a highway).
[0149] According to some embodiments, portion A-0 of FIG. 7G shows user 7320 with a line of sight directed at a given physical object (e.g., a hand in a ready state configuration) within a physical environment in a three-dimensional environment (e.g., the user's hand 7200 or its representation is within the field of view of one or more cameras of device 7100, or visible through a pass-through or transparent portion of a head-mounted display (HMD) or a heads-up display). In some embodiments, device 7100 uses one or more cameras (e.g., a front camera) directed at the user to track the movement of the user's eyes (or the movement of both eyes of the user) to determine the direction and / or target of the user's line of sight. Further details of exemplary line-of-sight tracking techniques are provided in some embodiments with respect to FIGS. 1-6, and particularly FIGS. 5 and 6. In portion A-0 of FIG. 7G, while hand 7200 is in a ready state, the user's line of sight is directed at hand 7200, so that (as indicated by the dotted line associating the user's eyeball 7512 with the representation of user 7200 or the user's hand 7200' (e.g., the actual hand or the representation of the hand presented via a display generation component)), system user interface operations (e.g., not user interface operations associated with other areas or elements of the user interface or with individual software applications running on device 7100, but system affordance 7214 or user interface operations associated with a system menu (e.g., menu 7170) associated with the system affordance) are executed in response to a gesture performed using hand 7200. Portion B-0 of FIG. 7G shows a user whose line of sight has moved away from a given physical object (e.g., a hand in a ready state or other given physical object) within a three-dimensional environment while the user's hand is in a ready state configuration (e.g., the user's hand 7200 or other given physical object, or its representation that is visible within the field of view of one or more cameras of device 7100 or through a pass-through or transparent portion of an HMD or a heads-up display).In some embodiments, the computer system requires that the user's line of sight remain on a predetermined physical object (e.g., the user's hand or another predetermined physical object) while the hand input gesture is being processed in order to provide a system response corresponding to the input gesture. In some embodiments, the computer system requires that the user's line of sight remain on a predetermined physical object for a threshold time with a predetermined stability (e.g., remaining substantially stationary or maintaining less than a threshold amount of movement for the threshold time) in order to provide a system response corresponding to the input gesture, and, optionally, after the time and stability requirements are met and before the input gesture is complete, the line of sight can move away from the predetermined input gesture.
[0150] In one example, portion A-1 of FIG. 7G shows a thumb flick gesture that begins in a hand ready state configuration and includes a forward movement of the thumb across the index finger of hand 7200 along the axis indicated by arrow 7120. In response to the thumb flick gesture starting from the ready state configuration of portion A-1 of FIG. 7G, according to a determination that the user's line of sight is directed at a predetermined physical object (e.g., hand 7200 as seen in the real world or via the display generation component of the computer system), the computer system displays a system menu 7170 (e.g., a menu of application icons) within the three-dimensional environment (e.g., replacing the system affordance 7214 with the tip of the user's thumb).
[0151] In another example, portion A-2 of FIG. 7G shows a thumb swipe gesture that starts from a ready state configuration and includes movement of the thumb of hand 7200 along the length of the index finger of hand 7200 along the axis indicated by arrow 7130. In this example, the hand gesture of portion A-2 of FIG. 7G is performed while the computer system is displaying a system menu (e.g., menu 7170) (e.g., system menu 7170 is displayed in response to the thumb flick gesture described herein with reference to portion A-1 of FIG. 7G). In response to the thumb swipe gesture of portion A-2 of FIG. 7G, starting from the ready state configuration and according to the determination that the user's line of sight is directed towards a predetermined physical object (e.g., hand 7200) (e.g., the line of sight meets predetermined position, duration, and stability requirements), the computer system moves the current selection indicator 7198 on the system menu (e.g., menu 7170) in the direction of movement of the thumb of hand 7200 (e.g., to an adjacent user interface object on menu 7170). In some embodiments, the thumb swipe gesture is one of a sequence of two or more input gestures that starts from the hand in the ready state configuration and represents a continuous series of user interactions with system user interface elements (e.g., system menus or system control objects, etc.). Thus, in some embodiments, the requirement for the line of sight to remain on a predetermined physical object (e.g., the user's hand) is optionally applied only to the start of the first input gesture (e.g., the thumb flick gesture of portion A-1 of FIG. 7G), and is not imposed on subsequent input gestures as long as the user's line of sight is directed towards the system user interface element during the subsequent input gestures. For example, according to the determination that the user's line of sight is directed towards the user's hand, or according to the determination that the user's line of sight is directed towards a system menu disposed at a position fixed relative to the user's hand (e.g., the tip of the thumb), the computer system performs a system operation in response to the thumb swipe gesture (e.g., navigate within the system menu).
[0152] In yet another example, portion A-3 of FIG. 7G shows a thumb tap gesture that starts from the hand in the ready state configuration and includes the movement of the thumb of hand 7200 that taps on the index finger of hand 7200 (e.g., the thumb moves from a raised position relative to the index finger and then moves downward from the raised position along the axis indicated by arrow 7110 until the thumb touches the index finger again). In this example, the thumb tap gesture of portion A-3 of FIG. 7G is executed while the computer system is displaying a system menu (e.g., menu 7170), and the current selection indicator is displayed on the corresponding user interface object (e.g., in response to the thumb swipe gesture described herein with reference to portion A-2 of FIG. 7G). In response to the thumb tap gesture of portion A-3 of FIG. 7G, in response to the user's line of sight being directed at a predetermined physical object (e.g., hand 7200), the currently selected user interface object 7190 is activated and an operation corresponding to the user interface object 7190 is executed (e.g., the display of menu 7170 is stopped, a user interface object 7306 (e.g., a preview or control panel, etc.) associated with the user interface object 7190 is displayed, and / or an application corresponding to the user interface object 7190 is launched). In some embodiments, the user interface object 7306 is displayed at a position within a three-dimensional environment corresponding to the position of hand 7200. In some embodiments, the thumb tap gesture is one of a sequence of two or more input gestures that starts from the hand in the ready state configuration and represents a continuous series of user interactions with system user interface elements (e.g., system menus or system control objects, etc.).Accordingly, in some embodiments, the requirement for the line of sight to remain on a predetermined physical object (e.g., the user's hand) is optionally applied only to the start of a first input gesture (e.g., the thumb flick gesture of portion A-1 in FIG. 7G), and not imposed on subsequent input gestures (e.g., the thumb swipe gesture of portion A-2 and the thumb tap gesture of portion A-3 in FIG. 7G) as long as the user's line of sight is directed at a system user interface element during the subsequent input gesture. For example, in accordance with a determination that the user's line of sight is directed at the user's hand, or a determination that the user's line of sight is directed at a system menu positioned at a fixed position relative to the user's hand (e.g., the tip of the thumb), the computer system performs a system operation in response to the thumb tap gesture (e.g., activates the currently selected user interface object within the system menu).
[0153] In contrast to portion A-0 of FIG. 7G, portion B-0 of FIG. 7G shows a user whose line of sight is away from a given physical object (e.g., a hand in a ready state configuration) within a three-dimensional environment (e.g., the user's hand 7200 or its representation is within the field of view of one or more cameras of device 7100, but the user's line of sight is not on the hand 7100 (e.g., directly, or via a pass-through or transparent portion of an HMD or head-up display, or via a camera view)). Instead, the user's line of sight is directed towards a container 7206 (e.g., a menu row) of a user interface object (e.g., as described herein with reference to container 7206 of FIGS. 7D - 7E) or the displayed user interface in general (e.g., the user's line of sight does not meet the stability and duration requirements for a particular location or object within the three-dimensional environment). In accordance with the determination that the user's line of sight is away from a given physical object (e.g., hand 7200), the computer system does not perform a system user interface operation (e.g., a system affordance 7214 or a user interface operation associated with a system menu (e.g., a menu of application icons) associated with a system affordance (e.g., as shown in portion A-1 of FIG. 7G)) in response to a thumb flick gesture performed using hand 7200. Optionally, instead of a system user interface operation, a user interface operation associated with another area or element of the user interface, or associated with an individual software application running on a computer system (e.g., device 100 or HMD), is performed in response to a thumb flick gesture performed using hand 7200 while the user's line of sight is directed away from a given physical object (e.g., a ready state hand 7200). In one example, as shown in FIG. 7G, the entire user interface including container 7206 is scrolled upward in accordance with an upward thumb flick gesture.In another example, the user interface operations associated with the container 7206 are performed in response to a gesture executed using the hand 7200 while the user's line of sight is directed from the hand 7200 to the container 7206, instead of system user interface operations (e.g., display of a system menu).
[0154] Similar to portion A-1 of FIG. 7G, portion B-1 of FIG. 7G also shows a thumb flick gesture by the hand 7200 starting from the ready state configuration. In contrast to the behavior shown in portion A-1 of FIG. 7G, in portion B-1 of FIG. 7G, according to the determination that the user's line of sight is not directed at a predetermined physical object (e.g., the user's hand in the ready state configuration), the computer system does not display the system menu in response to the thumb flick gesture. Instead, the user interface is scrolled upward according to the thumb flick gesture. For example, the container 7206 is moved upward in the three-dimensional environment according to the movement of the thumb of the hand 7200 across the index finger of the hand 7200 and according to the fact that the user's line of sight is directed from the hand 7200 to the container 7206. In this example, although the system operation is not executed, a system affordance 7214 indicating that the system operation is available (e.g., because the user's hand is in the ready state configuration) remains displayed next to the user's thumb, and the system affordance 7214 moves with the user's thumb while the user interface is scrolled upward in response to the input gesture during the input gesture.
[0155] Similar to portion A-2 of FIG. 7G, portion B-2 of FIG. 7G also shows a thumb swipe gesture by hand 7200 starting from the ready state configuration. In contrast to portion A-2 of FIG. 7G, the thumb swipe gesture of portion B-2 of FIG. 7G is executed while (for example, since menu 7170 is not displayed in response to the flick gesture described herein with reference to portion B-1 of 7G, and) the system menu (for example, menu 7170) is not displayed and the user's line of sight is not focused on a predetermined physical object (for example, the user's hand). In response to the thumb swipe gesture of portion B-2 of FIG. 7G, according to the determination that the user's line of sight is directed away from a predetermined physical object (for example, hand 7200) and towards container 7206, the current selection indicator within container 7206 is scrolled in the moving direction of the thumb of hand 7200.
[0156] Similar to portion A-3 of FIG. 7G, portion B-3 of FIG. 7G also shows a thumb tap gesture by hand 7200 starting from a ready state configuration. In contrast to portion A-3 of FIG. 7G, the thumb tap gesture of portion B-3 of FIG. 7G is executed while (for example, since menu 7170 is not displayed in response to a flick gesture described herein with reference to portion B-1 of 7G) the system menu (for example, menu 7170) is not displayed and the user's line of sight is not focused on a predetermined physical object (for example, the user's hand). In response to the thumb tap gesture of portion A-3 of FIG. 7G, in accordance with the fact that the user's line of sight is directed away from a predetermined physical object (for example, hand 7200, or in the case where there is no system user interface element in response to a previously received input gesture), the computer system does not perform the execution of system operations. In accordance with the determination that the user's line of sight is directed towards container 7206, the currently selected user interface object within container 7206 is activated, and the operation corresponding to the currently selected user interface object is executed (for example, the display of container 7206 is stopped, and user interface 7308 corresponding to the activated user interface object is displayed). In this example, although the system operation is not executed, the system affordance 7214 indicating that the system operation is available (for example, because the user's hand is in a ready state configuration) remains displayed next to the user's thumb, and the system affordance 7214 moves with the user's thumb while the user interface scrolls upward in response to the input gesture during the input gesture. In some embodiments, since the user interface object 7308 is activated from container 7206, in contrast to user interface object 7306 of portion A-3 of FIG. 7G, the user interface object 7308 is displayed at a position within a three-dimensional environment that does not correspond to the position of hand 7200.
[0157] In the example shown in FIG. 7G, it should be understood that the computer system treats the user's hand 7200 as a predetermined physical object whose position is used to determine whether system operations should be executed in response to a predetermined gesture input (e.g., compared to the user's line of sight). The position of the hand 7200 may appear different on the display in the examples shown in the left and right columns of FIG. 7G, which indicates that the position of the user's line of sight has changed with respect to the three-dimensional environment, and does not necessarily impose a limitation on the position of the user's hand with respect to the three-dimensional environment. In fact, in most situations, the entire hand of the user is often not fixed in position during the input gesture, and the line of sight is compared to the physical position where the user's hand moves to determine whether the line of sight is focused on the user's hand. In some embodiments, when another physical object in the user's environment other than the user's hand is used as a predetermined physical object to determine whether system operations should be executed in response to a predetermined gesture input, the line of sight is compared to the physical position of the physical object even if the physical object is moving with respect to the environment or the user, and / or the user is moving with respect to the physical object.
[0158] In the example shown in FIG. 7G, according to some embodiments, whether the user's line of sight is directed at a predetermined physical position (e.g., focused on a physical object (e.g., the user's hand in a ready state configuration or a system user interface object displayed at a fixed position relative to a predetermined physical object)) is used in conjunction with whether the input gesture is starting from a hand in a ready state configuration to determine whether to execute a system operation (e.g., display a system user interface or a system user interface object), or to execute an operation within the current context of the three-dimensional environment without executing the system operation. FIGS. 7H - 7J show exemplary behavior of a displayed three-dimensional environment (e.g., a virtual reality or mixed reality environment) that depends on whether the user's line of sight meets a predetermined requirement (e.g., the line of sight is focused on an activatable virtual object and meets stability and duration requirements), and whether the user is ready to provide a gesture input (e.g., whether the user's hand meets a predetermined requirement (e.g., has risen to a predetermined height and has been placed in a ready state configuration for at least a threshold time)). The input gestures described with reference to FIGS. 7H - 7J are used to exemplify the processes described below, including the process in FIG. 11.
[0159] FIG. 7H shows an exemplary computer-generated environment corresponding to a physical environment. As described herein with reference to FIG. 7H, the computer-generated environment can be a virtual reality environment, an augmented reality environment, or a computer-generated environment displayed on a display, and the computer-generated environment is displayed on the display so as to be superimposed on a view of the physical environment visible through a transparent portion of the display. As shown in FIG. 7H, user 7502 stands in a physical environment (e.g., scene 105) that operates a computer system (e.g., computer system 101) (e.g., holds device 7100 or wears an HMD). In some embodiments, as in the example shown in FIG. 7H, device 7100 is a handheld device (e.g., a mobile phone, a tablet, or other mobile electronic device) that includes a display, a touch-sensitive display, etc. In some embodiments, device 7100 represents a wearable headset that includes a head-up display, a head-mounted display, etc., and is optionally replaceable. In some embodiments, the physical environment includes one or more physical surfaces and physical objects (e.g., the walls of a room, represented by, e.g., the shadowed 3D box 7504) surrounding user 7502.
[0160] In the example shown in part (B) of FIG. 7H, a computer-generated three-dimensional environment corresponding to a physical environment (e.g., a portion of the physical environment that is within the field of view of one or more cameras of device 7100 or is visible through a transparent portion of the display of device 7100) is displayed on device 7100. The physical environment includes a physical object 7504 that is represented by an object 7504' in the computer-generated environment shown on the display (e.g., the computer-generated environment is a virtual reality environment that includes a virtual representation of the physical object 7504, the computer-generated environment is an augmented reality environment that includes a representation 7504' of the physical object 7504 as part of a live view of one or more cameras of device 7100, or the physical object 7504 is visible through a transparent portion of the display of device 7100). Further, the computer-generated environment shown on the display includes virtual objects 7506, 7508, and 7510. The virtual object 7508 is displayed so as to appear attached to the object 7504' (e.g., overlapping the flat front surface of the physical object 7504). The virtual object 7506 is displayed so as to appear attached to a wall of the computer-generated environment (e.g., overlapping a wall or a portion of the representation of a wall of the physical environment). The virtual object 7510 is displayed so as to appear attached to the floor of the computer-generated environment (e.g., overlapping the floor or a portion of the representation of the floor of the physical environment). In some embodiments, the virtual objects 7506, 7508, and 7510 are activatable user interface objects that, when activated by user input, perform object-specific actions. In some embodiments, the computer-generated environment also includes virtual objects that are not activatable by user input but are displayed to improve the aesthetic quality of the computer-generated environment and provide information to the user.According to some embodiments, portion (C) of FIG. 7H shows that the computer-generated environment shown on device 7100 is a three-dimensional environment, and as the perspective with respect to the physical environment of device 7100 changes (e.g., as the field of view angle with respect to the physical environment of device 7100 or one or more cameras of device 7100 changes in response to movement and / or rotation of device 7100 within the physical environment), accordingly, the perspective of the computer-generated environment displayed on device 7100 changes (e.g., including changes to physical surfaces and objects (e.g., walls, floors, physical object 7504) as well as virtual objects 7506, 7508, and 7510).
[0161] Figure 7I shows an exemplary behavior of a computer-generated environment in response to the user looking at each virtual object in the computer-generated environment while the user 7502 is not ready to provide a gesture input (e.g., the user's hand is not in a ready state configuration). As shown in parts (A)-(C) of Figure 7I, to provide a gesture input, the user keeps their left hand 7200 in a state other than the ready state (e.g., a position other than the first predetermined ready state configuration). In some embodiments, the computer system determines that the user's hand is in a predetermined ready state for providing a gesture input according to detecting that a predetermined portion of the user's finger touches a physical control element (e.g., the thumb touches the center of the index finger, or the index finger touches a physical controller). In some embodiments, the computer system determines that the user's hand is in a predetermined ready state for providing a gesture input according to detecting that the user's hand has risen to a predetermined height above the user (e.g., the hand has risen in response to rotation of the arm about the elbow joint, or rotation of the wrist about the wrist, or fingers raised relative to the hand). In some embodiments, the computer system determines that the user's hand is in a predetermined ready state for providing a gesture input according to detecting that the posture of the user's hand has changed to a predetermined configuration (e.g., the thumb is placed between the index finger, the fingers are closed to form a fist, etc.). In some embodiments, a combination of multiple of the above requirements is used to determine whether the user's hand is in a ready state for providing a gesture input. In some embodiments, the computer system also requires that the entire user's hand is stationary (e.g., less than a threshold movement amount without a threshold time) in order to determine that the hand is ready to provide a gesture input.According to some embodiments, if the user's hand is not found to be in a ready state for providing a gesture input and the user's line of sight is focused on a virtual object that can be activated, subsequent movement of the user's hand (e.g., free movement or movement mimicking a predetermined gesture) is not recognized and / or not treated as a user input directed to the virtual object that is the focus of the user's line of sight.
[0162] In this example, the representation of the hand 7200 is displayed in a computer-generated environment. The computer-generated environment does not include a representation of the user's right hand (e.g., because the right hand is not within the field of view of one or more cameras of the device 7100). Further, in some embodiments, for example, in the example shown in FIG. 7I where the device 7100 is a handheld device, the user can view a portion of the surrounding physical environment separately from any representation of the physical environment displayed on the device 7100. For example, a portion of the user's hand is visible to the user outside the display of the device 7100. In some embodiments, the device 7100 in these examples represents and can be replaced by a headset having a display (e.g., a head-mounted display) that completely blocks the user's view of the surrounding physical environment. In some such embodiments, no portion of the physical environment is directly visible to the user, and instead, the physical environment is made visible to the user through a representation of the portion of the physical environment displayed by the device. In some embodiments, while the current state of the user's hand is continuously or periodically monitored by the device to determine whether the user's hand(s) has entered a ready state for providing a gesture input, the user's hand(s) is / are not visible to the user directly or through the display of the device 7100. In some embodiments, the device displays an indicator of whether the user's hand is in a ready state for providing an input gesture, provides feedback to the user, and warns the user to adjust the position of the hand if the user desires to provide an input gesture.
[0163] In part (A) of FIG. 7I, the user's line of sight is directed towards the virtual object 7506 (as indicated, for example, by the dotted line connecting the representation of the user's eyeball 7512 and the virtual object 7506). In some embodiments, the device 7100 uses one or more cameras (e.g., a front camera) directed towards the user to track the movement of the user's eyes (or the movement of both of the user's eyes) in order to determine the direction and / or object of the user's line of sight. Further details of eye tracking or line of sight tracking techniques are provided in FIGS. 1-6, particularly FIGS. 5-6, and the accompanying description. In part (A) of FIG. 7I, in response to the user directing their line of sight towards the virtual object 7506, in accordance with the determination that the user's hand is not ready to provide a gesture input (e.g., the left hand is not stationary and has not been maintained in a first predetermined ready state configuration for more than a threshold time), no action is performed on the virtual object 7506. Similarly, in part (B) of FIG. 7I, the user's line of sight has the left virtual object 7506 and is currently directed towards the virtual object 7508 (as indicated, for example, by the dotted line connecting the representation of the user's eyeball 7512 and the virtual object 7508). In response to the user directing their line of sight towards the virtual object 7508, in accordance with the determination that the user's hand is not ready to provide a gesture input, no action is performed on the virtual object 7508. Similarly, in part (C) of FIG. 7I, the user's line of sight is directed towards the virtual object 7510 (as indicated, for example, by the dotted line connecting the representation of the user's eyeball 7512 and the virtual object 7510). In response to the user directing their line of sight towards the virtual object 7510, in accordance with the determination that the user's hand is not ready to provide a gesture input, no action is performed on the virtual object 7510.In some embodiments, in order to trigger a visual change indicating that a virtual object under the user's line of sight can be activated by a gesture input, it is required that the user's hand be in a ready state to provide the gesture input. This is not because the user wants to interact with a specific virtual object in the environment, but rather because the user simply wants to confirm the environment (for example, looking at various virtual objects for a short time or intentionally for a certain period). This is advantageous because it tends to prevent unnecessary visual changes in the display environment. As a result, when experiencing a computer-generated three-dimensional environment using a computer system, the user's visual fatigue, confusion, and thus the user's mistakes are reduced.
[0164] In contrast to the exemplary scenario shown in FIG. 7I, FIG. 7J shows an exemplary behavior of a computer-generated environment in response to the user looking at each virtual object in the computer-generated environment while the user is ready to provide a gesture input according to some embodiments. As shown in parts (A) - (C) of FIG. 7J, while virtual objects 7506, 7608, and 7510 are displayed in the three-dimensional environment, the user holds their hand in a first ready state configuration for providing a gesture input (for example, the thumb is placed on the index finger, and the hand is raised above the user's body at a predetermined height).
[0165] In part (A) of FIG. 7J, the user's line of sight is directed towards virtual object 7506 (as indicated, for example, by a dotted line connecting the representation of the user's eyeball 7512 and the virtual object 7506). In response to the user's line of sight being directed towards virtual object 7506 in accordance with the determination that the user's left hand is in a state ready to provide a gesture input (for example, the line of sight meets the requirements of duration and stability at the virtual object 7506), the computer system provides visual feedback indicating that the virtual object 7506 can be activated by a gesture input (for example, the virtual object 7506 is highlighted, enlarged, or expanded by additional information or user interface details to indicate that the virtual object 7506 is interactive (for example, one or more operations associated with the virtual object 7506 can be executed in response to the user's gesture input)). Similarly, in part (B) of FIG. 7J, the user's line of sight moves away from the virtual object 7506 and is now directed towards the virtual object 7508 (as indicated, for example, by a dotted line connecting the representation of the user's eyeball 7512 and the virtual object 7508). In response to the user's line of sight being directed towards the virtual object 7508 in accordance with the determination that the user's hand is in a state ready to provide a gesture input (for example, the line of sight meets the requirements of duration and stability at the virtual object 7508), the computer system provides visual feedback indicating that the virtual object 7506 can be activated by a gesture input (for example, the virtual object 7508 is highlighted, enlarged, or expanded by additional information or user interface details to indicate that the virtual object 7508 is interactive (for example, one or more operations associated with the virtual object 7508 can be executed in response to the user's gesture input)). Similarly, in part (C) of FIG. 7J, the user's line of sight moves away from the virtual object 7508 and is now directed towards the virtual object 7510 (as indicated, for example, by a dotted line connecting the representation of the user's eyeball 7512 and the virtual object 7510).In accordance with the determination that the user's hand is in a ready state to provide gesture input, in response to detecting that the user is looking at the virtual object 7510, the computer system provides visual feedback indicating that the virtual object 7506 can be activated by gesture input (e.g., the virtual object 7510 is highlighted, enlarged, or expanded by additional information or user interface details, indicating that the virtual object 7510 is interactive (e.g., indicating that one or more operations associated with the virtual object 7510 can be executed in response to gesture input)).
[0166] In some embodiments, while visual feedback indicating that a virtual object can be activated by gesture input is being displayed, and in response to detecting a gesture input starting from the user's hand in the ready state, the computer system performs an operation corresponding to the virtual object that is the target of the user's line of sight according to the user's gesture input. In some embodiments, the visual feedback indicating that each virtual object can be activated by gesture input stops being displayed in response to the user's line of sight moving away from each virtual object and / or the user's hand ceasing to be in the ready state to provide gesture input without providing a valid gesture input.
[0167] In some embodiments, each virtual object (e.g., virtual objects 7506, 7508, or 7510) corresponds to an application (e.g., each virtual object is an application icon), and the actions associated with each virtual object that are available to be executed include launching the corresponding application, executing one or more actions within the application, or displaying a menu of actions related to the application or executed within the application. For example, if each virtual object corresponds to a media player application, the one or more actions include increasing the output volume of the media (e.g., in response to a thumb swipe gesture or pinch and twist gesture in a first direction), decreasing the output volume (e.g., in response to a thumb swipe gesture or pinch and twist gesture in a second direction opposite the first direction), switching the playback of the media (e.g., play or pause, in response to a thumb tap gesture), fast-forwarding, rewinding, browsing the media for playback (e.g., in response to a plurality of consecutive thumb swipe gestures in the same direction), or otherwise controlling the media playback (e.g., menu navigation in response to a thumb flick gesture followed by a thumb swipe gesture). In some embodiments, each virtual object is a simplified user interface for controlling a physical object (e.g., an electronic device, smart speaker, smart lamp, etc.) underlying the corresponding virtual object, and upon detection of a wrist flick gesture or thumb flick gesture while a visual indication that the respective virtual objects are interactive is displayed, the computer system stops displaying an extended user interface for controlling the physical object (e.g., showing an on / off button, the currently playing media album, and additional playback controls and output adjustment controls, etc.).
[0168] In some embodiments, visual feedback indicating that a virtual object is interactive (e.g., in response to user input including other types of input such as gesture input, audio input, and touch input) includes displaying one or more user interface objects or information, prompts that were not displayed before the user's line of sight was directed onto the virtual object. In one example, each virtual object is a virtual window superimposed on a physical wall represented in a three-dimensional environment, and in response to the user directing their line of sight to the virtual window while their hand is in a ready state to provide gesture input, the computer system displays the position and / or date and time associated with the virtual landscape visible through the virtual window, indicating that the landscape can be changed (e.g., through changes in position, date and time, season, etc. according to subsequent gesture input by the user). In another example, each virtual object includes a displayed still photograph (e.g., each virtual object is a picture frame), and in response to the user directing their line of sight to the displayed photograph while their hand is in a ready state to provide gesture input, the computer system displays a multi-frame photograph or video clip associated with the displayed still photograph, indicating that the photograph is interactive, and optionally indicating that the photograph can be changed (e.g., by browsing a photo album according to subsequent gesture input by the user).
[0169] Figures 7K-7M show exemplary views that change in response to detection of a three-dimensional environment (e.g., a virtual reality environment or a mixed reality environment) that changes in response to detection of a user's hand grip on a housing of a display generation component of a computer system (e.g., computer system 101 of FIG. 1 (e.g., a handheld device or an HMD)) while the display generation component of the computer system (e.g., the display generation component 120 of FIGS. 1, 3, and 4) is positioned at a predetermined position relative to a user of the device (e.g., when the user first enters a computer-generated reality experience (e.g., when the user holds the device in front of their eyes or when the user wears an HMD on their head)). The change in the view of the three-dimensional environment is not fully determined by the computer system without user input, but rather is formed by the user (e.g., by changing their grip on the device or the housing of the display generation component) to initiate the first transition into the computer-generated reality experience. The input gesture described with reference to FIG. 7G is used to illustrate the processes described below, including the process in FIG. 12.
[0170] Part (A) of FIG. 7K shows the physical environment 7800 in which a user (e.g., user 7802) is using a computer system. The physical environment 7800 includes one or more physical surfaces (e.g., walls, floors, surfaces of physical objects, etc.) and physical objects (e.g., physical object 7504, the user's hand, body, etc.). Part (B) of FIG. 7K shows an exemplary view 7820 of a three-dimensional environment (also referred to as the "first view 7820" or the "first view 7820" of the three-dimensional environment) that is displayed by a display generation component of the computer system (e.g., device 7100 or HMD). In some embodiments, the first view 7820 is displayed when the display generation component (e.g., the display of device 7100 or HMD) is positioned at a predetermined position relative to the user 7802. For example, in FIG. 7K, the display of device 7100 is positioned in front of the user's eyes. In another example, the computer system determines that the display generation component is positioned at a predetermined position relative to the user according to the determination that the display generation component (e.g., HMD) is positioned on the user's head so that the user can view the physical environment only through the display generation component. In some embodiments, the computer system determines that the display generation component is positioned at a predetermined position relative to the user according to the determination that the user is seated in front of the head-up display of the computer system. In some embodiments, by positioning the display generation component at a predetermined position relative to the user, or positioning the user at a predetermined position relative to the display generation component, the user can view content (e.g., real or virtual content) through the display generation component. In some embodiments, when the display generation component and the user are in a predetermined relative position, the user's view of the physical environment can be at least partially (or completely) blocked by the display generation component.
[0171] In some embodiments, the placement of the display generation component of the computer system is determined based on the analysis of data captured by the sensor system. In some embodiments, the sensor system includes one or more sensors that are components of the computer system (e.g., internal components enclosed within the same housing as the display generation component of device 7100 or the HMD). In some embodiments, the sensor system is an external system and is not enclosed within the same housing as the display generation component of the computer system (e.g., the sensor is an external camera that provides image data captured for data analysis to the computer system).
[0172] In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras) that track the movement of the user of the computer system and / or the display generation component. In some embodiments, the one or more imaging sensors track the position and / or movement of one or more features of the user, such as the user's hand(s) and / or the user's head, to detect the placement of the display generation component relative to the user or a predetermined portion of the user (e.g., the head, eyes, etc.). For example, the image data is analyzed in real time to determine whether the user is holding the display of device 7100 in front of the user's eyes or whether the user is wearing a head-mounted display on the user's head. In some embodiments, one or more imaging sensors track the user's line of sight to determine where the user is looking (e.g., whether the user is looking at the display). In some embodiments, the sensor system includes one or more touch-based sensors (e.g., attached to the display) that detect the user's grip on the display, such as holding device 7100 with one or both hands and / or on the edge of the device, or holding a head-mounted display with both hands and placing the head-mounted device on the user's head. In some embodiments, the sensor system includes one or more motion sensors (e.g., accelerometers) and / or position sensors (e.g., gyroscopes, GPS sensors, and / or proximity sensors) that detect the movement and / or position information (e.g., position, height, and / or orientation) of the display of the electronic device to determine the placement of the display relative to the user. For example, the movement and / or position data is analyzed to determine whether the mobile device is lifted and facing the user's eyes or whether the head-mounted display is lifted and placed on the user's head. In some embodiments, the sensor system includes one or more infrared sensors that detect the positioning of a head-mounted display on the user's head. In some embodiments, the sensor system includes a combination of different types of sensors that provide data for determining the placement of the display generation component relative to the user.For example, the grip of the user's hand on the housing of the display generation component, the movement and / or orientation information of the display generation component, and the user's line-of-sight information are analyzed in combination to determine the placement of the display generation component relative to the user.
[0173] In some embodiments, based on the analysis of data captured by the sensor system, it is determined that the display of the electronic device is positioned at a predetermined position relative to the user. In some embodiments, the predetermined position of the display relative to the user indicates that the user is attempting to initiate a virtual and immersive experience using the computer system (e.g., start a 3D movie, enter a 3D virtual world, etc.). For example, the sensor data indicates that the user is holding the mobile device within both palms while the user's line of sight is directed at the display screen or while the user is holding and raising the head-mounted device at the user's head with both hands (e.g., the hand configuration shown in FIG. 7K). In some embodiments, the computer system enables the user to adjust the position of the display generation component relative to the user (e.g., fit comfortably, move the HMD so that the display is properly aligned with both eyes), and the changes in hand grip and position during this period do not trigger any change in the first view being displayed. In some embodiments, the first hand grip monitored for changes is not a grip for holding the display generation component, but a touch of the hand or finger on a specific part of the display generation component (e.g., a switch or control for turning on the HMD or starting the display of virtual content). In some embodiments, the combination of placing the HMD on the user's head with a hand grip and activating the control to start an immersive experience is the first hand grip monitored for changes.
[0174] In some embodiments, as shown in portion (B) of FIG. 7K, in response to detecting that the display is in a predetermined position relative to the user, a first view 7820 of the three-dimensional environment is displayed by a display generation component of the computer system. In some embodiments, the first view 7820 of the three-dimensional environment is a welcome / introduction user interface. In some embodiments, the first view 7820 includes a pass-through portion that includes a representation of at least a portion of the physical environment 7800 surrounding the user 7802.
[0175] In some embodiments, the pass-through portion is a transparent or translucent (e.g., see-through) portion of the display generation component that surrounds the user 7802's field of view and reveals at least a portion of the physical environment 7800 within the field of view. For example, the pass-through portion is part of a head-mounted display that is translucent (e.g., 50%, 40%, 30%, 20%, 15%, 10%, or less than 5% opacity) or transparent, and the user can view the real world surrounding the user through it without removing the display generation component. In some embodiments, the pass-through portion gradually transitions from translucent or transparent to completely opaque as the welding / introduction user interface changes to an immersive virtual or mixed reality environment, e.g., in response to a subsequent change in the user's hand grip indicating that the user is ready to proceed into a fully immersive environment.
[0176] In some embodiments, the pass-through portion of the first view 7820 displays a live feed of an image or video of at least a portion of the physical environment 7800 captured by one or more cameras (e.g., a rear camera (s) associated with a mobile device or a head-mounted display, or other cameras that supply image data to an electronic device). For example, the pass-through portion includes all or a portion of a display screen that displays a live image or video of the physical environment 7800. In some embodiments, one or more cameras are directed at a portion of the physical environment that is directly in front of the user's eyes (e.g., behind the display generation component). In some embodiments, one or more cameras are directed at a portion of the physical environment that is not directly in front of the user's eyes (e.g., in a different physical environment, or to the side or rear of the user).
[0177] In some embodiments, the first view 7820 of the three-dimensional environment includes three-dimensional virtual reality (VR) content. In some embodiments, the VR content includes one or more virtual objects corresponding to one or more physical objects (e.g., shelves and / or walls) within the physical environment 7800. For example, at least some of the virtual objects are displayed at positions within the virtual reality environment corresponding to the positions of the physical objects within the corresponding physical environment 7800 (e.g., the positions of the physical objects within the physical environment are determined using one or more cameras). In some embodiments, the VR content does not correspond to the physical environment 7800 viewable through the pass-through portion and / or is displayed independently of the physical objects within the pass-through portion. For example, the VR content includes virtual user interface elements (e.g., a virtual dock including user interface objects, or a virtual menu), or other virtual objects that are unrelated to the physical environment 7800.
[0178] In some embodiments, the first view 7820 of the three-dimensional environment includes three-dimensional augmented reality (AR) content. In some embodiments, one or more cameras (e.g., a rear camera (s) associated with a mobile device or a head-mounted display, or other camera that supplies image data to a computer system) continuously provide a live view of at least a portion of the surrounding physical environment 7800 within the field of view of the one or more cameras, and the AR content corresponds to that portion of the surrounding physical environment 7800 within the field of view of the one or more cameras. In some embodiments, the AR content at least partially includes the live view of the one or more cameras. In some embodiments, the AR content includes one or more virtual objects disposed in a portion of the live view (e.g., overlaid on a portion of the live view or appearing to block a portion of the live view). In some embodiments, the virtual objects are displayed at positions within the virtual environment 7820 that correspond to the positions of the corresponding objects within the physical environment 7800. For example, each virtual object is displayed in place of the corresponding physical object within the physical environment 7800 (e.g., overlaid on the view of the physical object, obscuring that view, and / or replacing that view).
[0179] In some embodiments, in a first view 7820 of the three-dimensional environment, a pass-through portion (e.g., representing at least a portion of the physical environment 7800) is surrounded by virtual content (e.g., VR and / or AR content). For example, the pass-through portion does not overlap with the virtual content on the display. In some embodiments, in the first view 7820 of the three-dimensional virtual environment, VR and / or AR virtual content is displayed instead of the pass-through portion (e.g., overlaid on or replacing the displayed content). For example, virtual content (e.g., a virtual dock or virtual start menu listing a plurality of virtual user interface elements) is overlaid on or occludes a portion of the physical environment 7800 that is displayed through a translucent or transparent pass-through portion. In some embodiments, the first view 7820 of the three-dimensional environment initially includes only the pass-through portion without any virtual content. For example, when the user first holds the device in the user's palm (e.g., as shown in FIG. 7L), or when the user first places the head-mounted display on the user's head, the user views the portion of the physical environment within the user's field of view or within the field of view of the live-feed camera through the see-through portion. Then, virtual content (e.g., a welcome / introduction user interface having virtual menus / icons) gradually fades in to overlay or occlude the pass-through portion over a period of time while the user's hand grip does not change. In some embodiments, the welcome / introduction user interface remains displayed (e.g., in a stable state having both virtual content and the physical world and the pass-through portion) as long as the user's hand grip does not change.
[0180] In some embodiments, by enabling a user's virtual immersion experience, the current view of the surrounding real-world user is temporarily blocked by a display generation component (e.g., by the presence of a display in front of the user's eyes and the muting function of a head-mounted display). This is done at a time point before the start of the user's virtual immersion experience. By having a pass-through portion within the welcome / introduction user interface, the transition from the visual recognition of the physical environment surrounding the user to the user's virtual immersion experience benefits from a more smoothly controlled transition (e.g., a cognitively gentle transition). This allows the user, rather than the computer system or content provider being instructed on the timing to transition to a fully immersive experience for all users, to have more control over how much time it takes to fully prepare for the immersive experience after viewing the welcome / introduction user interface.
[0181] FIG. 7M shows another exemplary view 7920 of a three-dimensional environment (also referred to as the "second view 7920 of the three-dimensional environment" or the "second view 7920") that is displayed by a display generation component of a computer system (e.g., on the display of device 7100). In some embodiments, the second view 7920 of the three-dimensional environment replaces the first view 7820 of the three-dimensional environment in response to detection of a change in the user's hand grip (e.g., a change from the hand configuration (e.g., two-handed grip) in part (B) of FIG. 7K to the hand configuration (e.g., one-handed grip) in part (B) of FIG. 7L) on the housing of a display generation component of a computer system that meets a first predetermined criterion (e.g., a criterion corresponding to detection of a sufficient reduction in user control or protection).
[0182] In some embodiments, changes in the user's hand(s) grip are detected by the sensor system as described above with reference to FIG. 7K. For example, one or more imaging sensors track the movement and / or position of the user's hand to detect changes in the user's hand grip. In another example, one or more touch-based sensors on the display detect changes in the user's hand grip.
[0183] In some embodiments, the first predetermined criterion for a change in the user's hand grip requires a change in the total number of hands detected on the display (e.g., from two hands to one hand, from one hand to no hands, or from two hands to no hands), a change in the total number of fingers in contact with the display generation component (e.g., from eight fingers to six fingers, from four fingers to two fingers, from two fingers to no fingers), a change from hand contact to no hand contact on the display generation component, a change in the contact position (e.g., from the palm to the fingers), and / or a change in the contact intensity on the display (e.g., resulting from a change in the hand posture, orientation, relative gripping force of different fingers on the display generation component, etc.). In some embodiments, a change in the user's hand grip on the display does not cause a change in a predetermined position of the display relative to the user (e.g., a head-mounted display remains on the user's head and covers the user's eyes). In some embodiments, a change in the hand grip indicates that the user is (e.g., gradually or intentionally) releasing the hand from the display and is ready to immerse in a virtual immersion experience.
[0184] In some embodiments, the first hand grip monitored for a change is not a grip for holding the display generation component, but a touch of the hand or finger to a specific part of the display generation component (e.g., a switch or control for turning on an HMD or starting the display of virtual content), and the first predetermined criterion for a change in the user's hand grip requires that the finger that touched a specific part of the display generation component (e.g., a finger that activates a switch or control for turning on an HMD or starting the display of virtual content) stops touching the specific part of the display generation component.
[0185] In some embodiments, a second view 7920 of the three-dimensional environment replaces at least a portion of the pass-through portion in the first view 7820 with virtual content. In some embodiments, the virtual content of the second view 7920 of the three-dimensional environment includes VR content (e.g., virtual object 7510 (e.g., a virtual user interface element or system affordance)), AR content (e.g., virtual object 7506 (e.g., a virtual window superimposed on a live view of a wall captured by one or more cameras)), and / or virtual object 7508 (e.g., a photograph or virtual control displayed instead of or superimposed on a part or all of the representation 7504' of the physical object 7504 in the physical environment).
[0186] In some embodiments, replacing the first view 7820 with the second view 7920 includes increasing the opacity of the pass-through portion (e.g., when the pass-through portion is implemented in a translucent or transparent state of the display), making the virtual content superimposed on the translucent or transparent portion of the display more visible and having higher color saturation. In some embodiments, the virtual content of the second view 7920 provides a more immersive experience for the user than the virtual content in the first view 7820. For example, the virtual content of the first view 7820 is displayed in front of the user, while the virtual content of the second view 7920 includes a three-dimensional world represented in a panorama or 360-degree view that is visible to the user as the user turns their head and / or walks around. In some embodiments, the second view 7920 includes a small pass-through portion that displays less or a smaller portion of the physical environment 7800 surrounding the user than the first view 7820. For example, the pass-through portion of the first view 7820 shows an actual window in one of the walls of the room where the user is located, and the pass-through portion of the second view 7920 shows a window in one of the walls replaced with a virtual window, resulting in a reduced area of the pass-through portion in the second view 7920.
[0187] Figure 7M shows yet another exemplary third view 7821 (e.g., a first view 7820 of a three-dimensional environment or a modified version thereof, or a different view) that is displayed by a display generation component of a computer system in response to detecting a re-establishment of an initial hand grip configuration on the housing of the display generation component (e.g., after a second view 7920 was displayed in response to detecting a required change in the hand grip as shown in Figure 7L). In some embodiments, the third view 7821 re-establishes a pass-through portion in response to detecting another change in the user's hand grip on the display generation component of the computer system (e.g., a change from the hand configuration of Figure 7L, or a change from no hand grip to the hand configuration of Figure 7M). In some embodiments, the change in the user's hand grip represents a re-establishment of the user's hand grip on the housing of the display generation component and indicates that the user desires to exit the virtual immersion experience (e.g., partially or fully, gradually or immediately).
[0188] In some embodiments, the sensor system detects a change in the total number of hands detected on the housing of the display generation component (e.g., from one hand to two hands, or from no hands to two hands), an increase in the total number of fingers contacting the housing of the display generation component, a change from no hand contact to hand contact on the housing of the display generation component, a change in the contact position (e.g., from a finger (single or plural) to the palm), and / or a change in the contact intensity on the housing of the display generation component. In some embodiments, re - establishing the user's hand grip changes the position and / or orientation of the display generation component (e.g., changing the position and angle of device 7100 relative to the environment of part (A) of FIG. 7M as compared to the angle of part (A) of FIG. 7L). In some embodiments, a change in the user's hand grip causes a change in the user's perspective relative to the physical environment 7800 (e.g., the field - of - view angle of device 7100 or one or more cameras of device 7100 relative to the physical environment 7800 changes). As a result, the perspective of the third view 7821 being displayed changes accordingly (e.g., including changing the perspective of the pass - through portion and / or virtual objects on the display).
[0189] In some embodiments, the pass-through portion of the third view 7821 is the same as the pass-through portion of the first view 7820 or at least increased relative to the pass-through portion (if any) within the second view 7920. In some embodiments, the pass-through portion of the third view 7821 shows a different perspective of the physical object 7504 within the physical environment 7800 as compared to the pass-through portion of the first view 7820. In some embodiments, the pass-through portion of the third view 7821 is the see-through portion of a display generation component that is transparent or translucent. In some embodiments, the pass-through portion of the third view 7821 displays a live feed from one or more cameras configured to capture image data of at least a portion of the physical environment 7800. In some embodiments, there is no virtual content displayed with the pass-through portion within the third view 7821. In some embodiments, the virtual content is paused or made translucent or reduced in color saturation within the third view 7821 and is displayed simultaneously with the pass-through portion within the third view 7821. When the third view is displayed, the user can resume the immersive experience by changing their hand grip again, as described with respect to FIGS. 7K - 7L.
[0190] FIGS. 7N - 7P show exemplary views of a three-dimensional virtual environment that changes in response to detection of a change in the user's position relative to objects (e.g., obstacles or targets) within the physical environment surrounding the user, according to some embodiments. The input gestures described with reference to FIGS. 7N - 7P are used to exemplify the processes described below, including the process in FIG. 13.
[0191] In portion (A) of FIG. 7N, user 7802 holds device 7100 within physical environment 7800. The physical environment includes one or more physical surfaces and physical objects (e.g., walls, floors, physical object 7602). Device 7100 displays virtual three-dimensional environment 7610 without displaying a pass-through portion showing the physical environment surrounding the user. In some embodiments, device 7100 can be replaced by a head-mounted display (HMD) or other computer system that includes a display generation component that blocks the user's view of the physical environment when displaying virtual environment 7610. In some embodiments, the HMD or display generation component of the computer system at least surrounds the user's eyes, and the user's view of the physical environment is partially or completely blocked by virtual content displayed by the display generation component and other physical barriers formed by the display generation component or its housing.
[0192] Portion (B) of FIG. 7N shows a first view 7610 of a three-dimensional environment displayed by a display generation component (also referred to as a “display”) of a computer system (e.g., device 7100 or HMD).
[0193] In some embodiments, the first view 7610 is a three-dimensional virtual environment that provides an immersive virtual experience (e.g., a three-dimensional movie or game). In some embodiments, the first view 7610 includes three-dimensional virtual reality (VR) content. In some embodiments, the VR content includes one or more virtual objects corresponding to one or more physical objects in a physical environment that does not correspond to physical environment 7800 surrounding the user. For example, at least some of the virtual objects are displayed at positions within the virtual reality environment corresponding to the positions of physical objects in a physical environment remote from physical environment 7800. In some embodiments, the first view includes virtual user interface elements (e.g., a virtual dock or virtual menu including user interface objects), or other virtual objects that are unrelated to physical environment 7800.
[0194] In some embodiments, the first view 7610 does not include any representation of the physical environment 7800 surrounding the user 7802 and is instead 100% virtual content (e.g., virtual objects 7612 and virtual surfaces 7614 (e.g., virtual walls and floors)). In some embodiments, the virtual content (e.g., virtual objects 7612 and virtual surfaces 7614) within the first view 7610 does not correspond to or visually convey the presence, location, and / or physical structure of any physical object within the physical environment 7800. In some embodiments, the first view 7610 optionally includes a virtual representation that indicates the presence and location of a first physical object within the physical environment 7800 but does not visually convey the presence, location, and / or physical structure of a second physical object within the physical environment 7800, both of which are within the user's field of view if the user's view is not blocked by the display generation component. In other words, the first view 7610 includes virtual content that replaces the display of at least some physical objects or portions thereof that would be present in the user's normal field of view (e.g., the user's field of view with no display generation component placed in front of the user's eyes).
[0195] FIG. 7O shows another exemplary view 7620 of a three-dimensional virtual environment (also referred to as the "second view 7620 of the three-dimensional environment", the "second view 7620 of the virtual environment", or the "second view 7620") that is displayed by a display generation component of a computer system. In some embodiments, the sensor system detects that user 7802 is moving towards physical object 7602 within physical environment 7800, and the sensor data acquired by the sensor system is analyzed to determine whether the distance between user 7802 and physical object 7602 is within a predetermined threshold distance (e.g., the length of an arm, or the length of a user's normal step). In some embodiments, when it is determined that a portion of physical object 7602 is within the threshold distance with respect to user 7602, the appearance of the view of the virtual environment is changed to show some of the physical characteristics of the portion of physical object 7602 (e.g., showing portion 7604 of physical object 7602 within the second view 7620 of portion (B) of FIG. 7O that is within the threshold distance from user 7802, and not showing other portions of physical object 7602 that are within the same field of view of the user but not within the user's threshold distance). In some embodiments, instead of replacing a portion of the virtual content with a direct view or a camera view of portion 7604 of the physical object, the visual characteristics (e.g., opacity, color, texture, virtual material, etc.) of a portion of the virtual content at the location corresponding to portion 7604 of the physical object are changed to show the physical characteristics (e.g., size, color, pattern, structure, contour, shape, surface, etc.) of portion 7604 of the physical object. The change to the virtual content at the location corresponding to portion 7604 of physical object 7602 does not apply to other portions of the virtual content, including the portion of the virtual content at the location corresponding to the portion of physical object 7602 outside of portion 7604. In some embodiments, the computer system provides a blend between the portion of the virtual content at the location corresponding to portion 7604 of the physical object and the portion of the virtual content immediately outside of the location corresponding to portion 7604 of the physical object (e.g., to smooth the visual transition).
[0196] In some embodiments, the physical object 7602 is a static object within a physical environment 7800 such as a wall, chair, or table. In some embodiments, the physical object 7602 is an object that moves within the physical environment 7800, for example, another person or a dog within the physical environment 7800 that moves relative to the user 7802 while the user 7802 remains stationary relative to the physical environment 7800 (e.g., while the user is sitting on a sofa watching a movie, the user's pet moves around).
[0197] In some embodiments, while the user 7802 is enjoying a three-dimensional immersive virtual experience (including, for example, a panoramic three-dimensional display having surround sound effects and other virtual perceptions), real-time analysis of sensor data from a sensor system coupled to the computer system indicates that the user 7802 is close enough to the physical object 7602 (e.g., either by the movement of the user towards the physical object or the movement of the physical object towards the user), the user 7802 can benefit from receiving a warning to smoothly and unobtrusively blend with the virtual environment. This allows the user to make a more informed decision about changing their movement and / or pausing / resuming the immersive experience without losing the immersive quality of the experience.
[0198] In some embodiments, the second view 7620 is displayed when analysis of the sensor data indicates that the user 7802 is within a threshold distance of at least a portion of the physical object 7602 within the physical environment 7800 (e.g., the physical object 7602 has a range that can potentially be visible to the user based on the user's field of view with respect to the virtual environment). In some embodiments, assuming the position of a portion of the physical object relative to the user within the physical environment 7800, that portion of the physical object would be visible within the user's field of view if the display had a pass-through portion or if the display generation component were not present in front of the user's eyes.
[0199] In some embodiments, the portion 7604 within the second view 7620 of the virtual environment includes a translucent visual representation of the corresponding portion of the physical object 7602. For example, the translucent representation overlays the virtual content. In some embodiments, the portion 7604 of the second view 7620 of the virtual environment includes a glassy appearance of the corresponding portion of the physical object 7602. For example, as the user 7802 approaches a table placed in a room while enjoying an immersive virtual experience, the portion of the table closest to the user is shown with a glassy translucent see-through appearance that overlays the virtual content (e.g., a virtual ball or virtual grass in the virtual view), and the virtual content behind that portion of the table is visible through that portion of the glassy table. In some embodiments, the second view 7620 of the virtual environment exhibits a predetermined distortion or other visual effect (e.g., visual effects such as shimmering, undulating, glowing, dimming, blurring, swirling, or different text effects) applied to the portion 7604 corresponding to the portion of the physical object 7602 closest to the user 7802.
[0200] In some embodiments, when the user moves towards the corresponding part of the physical object 7602 and enters within the threshold distance, the second view 7620 of the virtual environment instantaneously replaces the first view 7610 so as to provide the user with a timely warning. In some embodiments, the second view 7620 of the virtual environment is gradually displayed, for example, using a fade-in / fade-out effect, providing a smoother transition and a less confusing / distracting user experience. In some embodiments, the computer system enables the user to navigate within the three-dimensional environment by moving within the physical environment and changes the view of the three-dimensional environment presented to the user to reflect the computer-generated movement within the three-dimensional environment. For example, as shown in FIGS. 7N and 7O, when the user is walking towards the physical object 7602, the user perceives their movement as moving in the same direction towards the virtual object 7612 within the three-dimensional virtual environment (e.g., the virtual object 7612 gets closer and larger). In some embodiments, the virtual content presented to the user is independent of the user's movement within the physical environment and does not change in accordance with the user's movement within the physical environment, except when the user reaches within the threshold distance of a physical object within the physical environment.
[0201] FIG. 7P shows yet another exemplary view 7630 of the three-dimensional environment (also referred to as the "third view 7630 of the three-dimensional environment", the "third view 7630 of the virtual environment", or the "third view 7630") displayed by the display generation component (e.g., device 7100 or HMD) of the computer system. In some embodiments, after displaying the second view 7620 of the virtual three-dimensional environment as described with reference to FIG. 7O, as the user 7802 continues to move towards the physical object 7602 within the physical environment 7800, the analysis of the sensor data indicates that the distance between the user 7802 and the portion 7606 of the physical object 7602 is less than a predetermined threshold distance. In response, the display transitions from the second view 7620 to the third view 7630. In some embodiments, depending on the structure (e.g., size, shape, length, width, breadth, etc.) and the relative position of the user and the physical object 7602, the portions 7606 and 7604 of the physical object 7602 that were within a predetermined threshold distance of the user when the user was at various positions within the physical environment 7800 are optionally completely distinct non-overlapping portions of the physical object, portion 7606 optionally completely surrounds portion 7604, portions 7606 and 7604 optionally partially overlap, or portion 7604 optionally completely surrounds portion 7606. In some embodiments, portions or the whole of one or more other physical objects can be visually represented, or stop being represented, within the currently displayed view of the virtual three-dimensional environment as the user moves around the room with respect to those physical objects, depending on whether those physical objects are inside or outside the user's predetermined threshold distance.
[0202] In some embodiments, the computer system optionally enables a user to pre-select a subset of physical objects within a physical environment 7800 where the distance between the user and a given physical object is monitored and visual changes are applied to the virtual environment. For example, the user can pre-select furniture and pets as a subset of physical objects and choose not to select items such as clothing or curtains as a subset of physical objects, and not apply visual changes to the virtual environment to warn the user about the presence of clothing or curtains even when the user is walking. In some embodiments, the computer system applies visual effects (e.g., transparency, opacity, shine, refractive index, etc.) to a portion of the virtual environment corresponding to each position of the physical object by the user, regardless of whether the user is within a threshold distance of the physical object, thereby always enabling one or more physical objects to be pre-specified as being visually represented in the virtual environment. These visual indications enable the user to orient themselves with respect to the real world and feel a safer and more stable sense when exploring the virtual world, even when the user is immersed in the virtual world.
[0203] In some embodiments, as shown in FIG. 7P, the third view 7630 includes a rendering of a portion 7606 of the physical object 7602 that is within a threshold distance from the user when the user approaches the physical object. In some embodiments, the computer system optionally further increases the value of the display characteristics of the visual effects applied to the portion of the virtual environment that shows the physical characteristics of the corresponding portion of the physical object 7602 according to the reduced distance between the user and a part of the physical object. For example, the computer system optionally increases the refractive index, color saturation, visual effects, opacity, and / or transparency of a part of the virtual environment corresponding to the portion of the physical object in the third view 7630 as the user gradually approaches that portion of the physical object. In some embodiments, as the user 7802 approaches the physical object 7602, the spatial extent of the visual effect increases and the corresponding portion of the physical object 7602 appears larger in the user's field of view with respect to the virtual environment. For example, as the user 7802 approaches the physical object 7602 in the physical environment 7800, the portion 7606 of the third view 7630 gradually increases in size from the virtual object 7612 and appears to extend towards the user for at least two reasons: (1) more portions of the physical object 7602 come within a predetermined distance of the user, and (2) the same portion of the physical object 7602 (e.g., portion 7604) occupies a larger portion of the user's field of view of the virtual environment as it approaches the user's eye, as compared to the portion 7604 of the second view 7620.
[0204] In some embodiments, the computer system defines a gesture input (e.g., the user raises one or both arms to a predetermined height relative to the user's body within a threshold time (e.g., a quick and sudden movement that is a muscle reflex to prevent falling or colliding with something)) that visually displays within the virtual environment by modifying the display characteristics of the virtual environment at the positions corresponding to those portions of the physical object that are partially located within the threshold distance of the user (e.g., all portions that are potentially visible within the user's field of view in the virtual environment), or all physical objects that are potentially visible within the user's field of view in the virtual environment. This feature helps enable the user to quickly reorient themselves without completely disengaging from the immersive experience when the user is confident of their body's position within the physical environment.
[0205] Additional explanations regarding FIGS. 7A - 7P are provided below with reference to methods 8000, 9000, 10000, 11000, 12000, and 13000 described with respect to FIGS. 8 - 13 below.
[0206] FIG. 8 is a flowchart of an exemplary method 8000 for interacting with a three-dimensional environment using a predetermined input gesture according to some embodiments. In some embodiments, method 8000 is executed in a computer system (e.g., computer system 101 of FIG. 1) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera facing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth sensing cameras)) or a camera facing forward from the user's head). In some embodiments, method 8000 is stored on a non-transitory computer-readable storage medium and is executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some operations of method 8000 are optionally combined and / or the order of some operations is optionally changed.
[0207] In method 8000, a computer system displays a view of a three-dimensional environment (e.g., a virtual or mixed reality environment) (8002). While displaying the view of the three-dimensional environment, the computer system uses one or more cameras (e.g., as opposed to using a touch-sensing glove, or a touch-sensing surface on an input device controlled by hand, or other non-image-based means such as sound waves, uses one or more cameras positioned at the lower edge of the HMD) to detect movement of the user's thumb on the user's index finger of the user's first hand (e.g., the left or right hand not wearing a glove or not covered or attached to the input device / surface). This is shown, for example, in FIG. 7A and the accompanying description (e.g., thumb tap, thumb swipe, and thumb flick gestures). In some embodiments, the user's hand or its graphical representation is displayed in the view of the three-dimensional environment (e.g., in a pass-through portion of the display generation component or as part of an augmented reality view of the physical environment surrounding the user). In some embodiments, the user's hand or its graphical representation is not shown in the view of the three-dimensional environment or is displayed on a portion of the display outside the view of the three-dimensional environment (e.g., a separate or floating window). The benefit of using one or more cameras, particularly cameras that are part of the HMD, includes that the spatial position and size of the user's hand are visually perceived by the user as if they naturally exist within the physical or virtual environment with which the user is interacting, without the need for additional calculations to match the space of the input device to the three-dimensional environment and / or to scale, rotate, and translate the representation of the user's hand before placing it in the displayed three-dimensional environment, giving the user an intuitive sense of scale, orientation, and fixed position for perceiving the three-dimensional environment on the display.Referring back to FIG. 8, in response to detecting (8006) the movement of the user's thumb on the user's index finger using one or more cameras (e.g., as opposed to more exaggerated gestures such as waving a finger or hand in the air or sliding it across a touch-sensitive surface), and the movement being a swipe of the thumb on the index finger of the first hand in a first direction (e.g., movement along a first axis (e.g., the x-axis) of the x and y axes, where movement along the x-axis is movement along the length of the index finger and movement along the y-axis is movement in a direction across the index finger (substantially perpendicular to movement along the length of the index finger)), the computer system performs a first action (e.g., changing a selected user interface object within the displayed user interface (e.g., repeating the selection of items in a first direction through a list of items corresponding to the first direction (in a left-right column of items)), adjusting the position of a user interface object within the displayed user interface (e.g., moving the object in a direction corresponding to the first direction (e.g., left or right) within the user interface), and / or adjusting the system settings of the device (e.g., adjusting the volume, moving to the next list item, moving to the previous list item, skipping forward (e.g., fast-forwarding and / or advancing to the next chapter, audio track, and / or content item), skipping backward (e.g., rewinding and / or moving to the previous chapter, audio track, and / or content item))). In some embodiments, a swipe in a first sub-direction of the first direction (e.g., towards the tip of the index finger) (e.g., along the length of the index finger) corresponds to performing the first action in one way, and a swipe in a second sub-direction of the first direction (e.g., towards the base of the index finger) corresponds to performing the first action in another way. This is shown, for example, in FIGS. 7B, 7C, and 7F, and the accompanying description.Referring back to FIG. 8, in response to detecting (8006) the movement of a user's thumb by the user's index finger using one or more cameras (e.g., a movement that is contrasted with more dramatic gestures such as waving a finger or hand in the air or sliding it across a touch-sensitive surface), and according to a determination that the movement is a tap of the thumb on the index finger at a first position on the index finger (e.g., in a first portion of the index finger such as the distal phalanx, middle phalanx, and / or proximal phalanx) that includes a touch down and lift off of the thumb on the index finger within a threshold time, the computer system performs a second operation that is different from the first operation (e.g., performs an operation corresponding to the currently selected user interface object and / or changes the selected user interface object within the displayed user interface). In some embodiments, performing the first / second operation includes changing the view of the three-dimensional user interface, and this change depends on the current operating context. In other words, each gesture triggers different actions depending on the current operating context (e.g., the object the user is looking at, the direction the user is facing, the last function performed immediately prior to the current gesture, and / or the currently selected object), and accordingly changes within the view of the three-dimensional environment in respective ways. This is shown, for example, in FIGS. 7B, 7C, and 7F, and the accompanying description.
[0208] In some embodiments, in response to detecting movement of a user's thumb on the user's index finger using one or more cameras, according to a determination that the movement is a swipe of the thumb on the index finger of the first hand in a second direction substantially perpendicular to the first direction (e.g., the second axis of the x-axis and y-axis (e.g., the y-axis), movement along the x-axis is movement along the length of the index finger, and movement along the y-axis is movement in a direction across the index finger (substantially orthogonal to movement along the length of the index finger)), the computer system performs a third operation different from the first and second operations (e.g., changes a selected user interface object in the displayed user interface (e.g., repeats selection in the second direction within a list of items corresponding to the second direction (e.g., up and down or vertically in a list of multiple rows of items in a 2D menu), adjusts the position of the user interface object in the displayed user interface (moves the object up and down or in the direction corresponding to the second direction within the user interface), and / or adjusts the system settings (e.g., volume) of the device). In some embodiments, the third operation is different from the first operation and / or the second operation. In some embodiments, a swipe in a first sub-direction of the second direction (e.g., away from the palm along the circumference of the index finger) corresponds to performing the third operation in one way, and a swipe in a second sub-direction of the second direction (e.g., towards the palm along the circumference of the index finger) corresponds to performing the third operation in another way.
[0209] In some embodiments, in response to detecting movement of the user's thumb on the user's index finger using one or more cameras, the computer system, in accordance with a determination that the movement is movement of the thumb on the index finger in a third direction (and a second direction) different from the first direction (and not a tap of the thumb on the index finger), performs a fourth operation different from the first operation, different from the second operation, and (and different from the third operation). In some embodiments, the third direction is an upward direction from the index finger away from the index finger (e.g., opposite to tapping on the side of the index finger), and the gesture is a flick of the thumb from the side of the index finger away from the index finger and palm. In some embodiments, this upward flick gesture across the middle of the index finger using the thumb pushes the currently selected user interface object into a three-dimensional environment and initiates an immersive experience (e.g., a 3D video or 3D virtual experience, a panoramic display mode, etc.) corresponding to the currently selected user interface object (e.g., a video icon, an app icon, an image, etc.). In some embodiments, while the immersive experience is ongoing, swiping downward across the middle of the index finger towards the palm (e.g., a movement in one of the sub-directions of the second direction, as opposed to tapping in the middle of the index finger) pauses, stops, and / or reduces the immersive state of the immersive experience (e.g., non-full screen, 2D mode, etc.).
[0210] In some embodiments, performing the first operation includes increasing a value corresponding to the first operation (e.g., a value of a system setting, a value indicating a position and / or selection of at least a part of a user interface (e.g., a user interface object), and / or a value corresponding to selected content or a part of the content). For example, increasing the value includes increasing the volume according to a determination that a swipe of the thumb on the index finger in a first direction (e.g., a direction along the length of the index finger or a direction along the circumference of the index finger) moves towards a first predetermined portion of the index finger (e.g., towards the tip of the index finger or towards the back side of the index finger), moving an object in an increasing direction (e.g., upward and / or rightward), and / or adjusting a position to a next or another forward position (e.g., within a list and / or content item). In some embodiments, performing the first operation further includes decreasing the value corresponding to the first operation according to a determination that a swipe of the thumb on the index finger in a first direction (e.g., a direction along the length of the index finger or a direction along the circumference of the index finger) moves away from a first predetermined portion of the index finger (e.g., away from the tip of the index finger or away from the back side (on the back side of the hand) of the index finger) towards a second predetermined portion of the index finger (e.g., towards the base of the index finger or towards the front side (on the palm side of the hand) of the index finger) (e.g., decreasing the value includes decreasing the volume, moving an object in a decreasing direction (e.g., downward and / or leftward), and / or adjusting a position to a previous or another preceding position (e.g., within a list and / or content item)). In some embodiments, the direction in which the thumb is swiped on the index finger in a second direction also determines the direction of a third operation in the same way as the direction of the swipe in the first direction determines the direction of the first operation.
[0211] In some embodiments, performing the first operation includes adjusting a value corresponding to the first operation (e.g., a value of a system setting, a value indicating a position and / or selection of at least a part of a user interface (e.g., a user interface object), and / or a value corresponding to selected content or a part of the content) by an amount corresponding to the amount of movement of the thumb on the index finger. In some embodiments, the movement of the thumb is measured relative to a threshold position on the index finger, and the value corresponding to the first operation is adjusted among a plurality of individual altitudes according to which threshold position is reached. In some embodiments, the movement of the thumb is measured continuously, and the value corresponding to the first operation is adjusted continuously and dynamically based on the current position of the thumb on the index finger (e.g., along or around the index finger). In some embodiments, the speed of movement of the thumb is used to determine the magnitude of the operation and / or the thresholds used to determine when different discrete values of the operation are triggered.
[0212] In some embodiments, in response to detecting movement of a user's thumb on the user's index finger using one or more cameras, the computer system, according to a determination that the movement is a tap of the thumb on the index finger at a second position different from a first position on the index finger (e.g., a portion of the index finger and / or a finger doubling), performs a fifth operation different from a second operation (e.g., performs an operation corresponding to the currently selected user interface object and / or changes the selected user interface object within the displayed user interface). In some embodiments, the fifth operation is different from the first operation, the third operation, and / or the fourth operation. In some embodiments, tapping the middle portion of the index finger activates the currently selected object, and tapping the tip of the index finger minimizes / halts / closes the currently active application or experience. In some embodiments, detecting a tap of the thumb on the index finger does not require detecting a lift-off of the thumb from the index finger, and while the thumb remains on the index finger, movement of the thumb or the entire hand can be treated as a movement combined with a tap-and-hold input of the thumb, for example, to drag an object.
[0213] In some embodiments, the computer system uses one or more cameras to detect a swipe of the user's thumb via the user's middle finger (e.g., while detecting the user's index finger extended away from the middle finger). In response to detecting a swipe of the user's thumb on the user's middle finger, the computer system performs a sixth operation. In some embodiments, the sixth operation is different from the first operation, the second operation, the third operation, the fourth operation, and / or the fifth operation. In some embodiments, the swipe of the user's thumb on the middle finger includes movement of the thumb along the length of the middle finger (e.g., movement from the base towards the tip of the middle finger, or vice versa), and one or more different operations are performed according to a determination that the swipe of the user's thumb on the middle finger includes movement of the thumb from the tip of the middle finger towards the base of the middle finger along the length of the middle finger, and / or movement of the thumb from the palmar side of the middle finger across the middle finger to the upper part of the middle finger.
[0214] In some embodiments, the computer system uses one or more cameras to detect a tap of the user's thumb on the user's middle finger (e.g., while detecting the user's index finger extended away from the middle finger). In response to detecting a tap of the user's thumb on the user's middle finger, the computer system performs a seventh operation. In some embodiments, the seventh operation is different from the first operation, the second operation, the third operation, the fourth operation, the fifth operation, and / or the sixth operation. In some embodiments, different operations are performed according to a determination that the tap of the user's thumb on the middle finger is at a first position on the middle finger and that the tap of the user's thumb on the middle finger is at a second position different from the first position on the middle finger. In some embodiments, an upward flick from the first and / or second position on the middle finger causes the device to perform other operations different from the first, second,..., and / or seventh operations.
[0215] In some embodiments, the computer system displays a visual indication of the operating context of a thumb gesture (e.g., a swipe / tap / flick on another finger of the thumb hand) within a three-dimensional environment (e.g., a menu of selectable options, a dial for adjusting a value, an avatar of a digital assistant, a selection indicator for the currently selected object, highlighting of an interaction object, etc.). For example, when the device detects that the user's hand is in or has entered a predetermined ready state (e.g., the thumb is placed on the side of the index finger or floating above the side of the index finger, and / or the back of the thumb is facing upward / placed on the side of the index finger and the wrist is flicked), and / or when the thumb side of the hand is facing upward towards the camera, the device displays a plurality of user interface objects within the three-dimensional environment, and the user interface objects respond to swipe and tap gestures on another finger of the thumb hand. Performing a first action (or a second, third, etc. action) includes displaying a visual change in the three-dimensional environment corresponding to the performance of the first action (or a second, third, etc. action). For example, displaying a visual change includes activating each of the plurality of user interface objects and causing the actions associated with each user interface object to be performed.
[0216] In some embodiments, while displaying a visual indication of the operating context of a thumb gesture (e.g., while displaying a plurality of user interface objects within a three-dimensional environment in response to detecting that the user's hand is in a predetermined ready state), the computer system uses one or more cameras to detect movement of the user's first hand (e.g., movement of the entire hand within the physical environment relative to the camera as opposed to relative internal movement of the fingers). For example, the computer system detects movement of the hand while the hand remains in the ready state (e.g., detects movement and / or rotation of the hand / wrist within the three-dimensional environment). In response to detecting movement of the first hand, the computer system changes the displayed position of the visual indication of the operating context of the thumb gesture (e.g., a plurality of user interface objects) within the three-dimensional environment according to the detected change in the position of the hand. For example, the computer system maintains the display of a plurality of user interface objects within a predetermined distance of the hand (e.g., during movement of the hand, the object menu sticks to the tip of the thumb). In some embodiments, the visual indication is a system affordance (e.g., an indicator regarding an application launch user interface or a dock). In some embodiments, the visual indication is a dock that includes a plurality of application launch icons. In some embodiments, the dock changes as a function of changes in the hand configuration (e.g., position of the thumb, position of the index / middle finger). In some embodiments, the visual indication disappears when the hand moves out of the micro gesture orientation (e.g., raises the thumb below the shoulder). In some embodiments, the visual indication reappears when the hand moves back into the micro gesture orientation. In some embodiments, the visual indication appears in response to a gesture (e.g., the user swipes the thumb over the index finger while looking at the hand). In some embodiments, the visual indication is reset (e.g., disappears) after a hand inactivity time threshold (e.g., 8 seconds). Further details are described, for example, with respect to FIGS. 7D-7F and 9, and the accompanying description.
[0217] In some embodiments, the computer system uses one or more cameras to detect movement of the user's thumb on the user's index finger of a second hand (e.g., different from the first hand), while (e.g., in a two-handed gesture scenario) detecting movement of the user's thumb on the user's index finger of the first hand, or (e.g., in a one-handed gesture scenario) without detecting movement of the user's thumb on the user's index finger of the first hand. In response to detecting movement of the user's thumb on the index finger of the second hand using one or more cameras, according to a determination that the movement is a swipe of the thumb on the index finger of the second hand in a first direction (e.g., along the length of the index finger, or around the index finger, or away from the side of the index finger upwards), the computer system performs an eighth operation different from the first operation, and according to a determination that the movement is a tap of the thumb on the index finger of the second hand at a first position (e.g., in a first portion of the index finger such as the distal phalanx, middle phalanx, and / or proximal phalanx), on the index finger of the second hand, the computer system performs a ninth operation different from the second operation (and the eighth operation). In some embodiments, the eighth and / or ninth operations are different from the first operation, the second operation, the third operation, the fourth operation, the fifth operation, the sixth operation, and / or the seventh operation. In some embodiments, when both hands are used to perform a two-handed gesture, the movement of the thumbs of both hands is processed as simultaneous input and used together to determine which function is triggered. For example, when the thumbs move from the tips of the index fingers of both hands (the hands are facing each other) towards the bases of the index fingers, the device expands the currently selected object, and when the thumbs move from the bases of the index fingers of both hands (the hands are facing each other) towards the tips of the index fingers, the device minimizes the currently selected object.In some embodiments, when the thumb taps the index fingers of both hands simultaneously, the device activates the currently selected object in a first way (e.g., starts video recording using the camera app), when the thumb taps the index finger of the left hand, the device activates the currently selected object in a second way (e.g., performs autofocus using the camera app), and when the thumb taps the index finger of the right hand, the device activates the currently selected object in a third way (e.g., takes a snapshot using the camera app).
[0218] In some embodiments, in response to detecting, using one or more cameras, movement of the user's thumb on the user's index finger (as opposed to more exaggerated gestures such as waving a finger or hand in the air or sliding on a touch-sensitive surface), if the movement is determined to include a thumb touch-down on the index finger of the first hand followed by a first hand wrist flick gesture (e.g., upward movement of the first hand relative to the first hand's wrist), the computer system performs a tenth action that is different from the first action (e.g., different from each or a subset of the first through ninth actions corresponding to other types of movement patterns of the user's finger), such as providing an input to operate a selected user interface object, providing an input to select an object (e.g., a virtual object selected and / or held by the user), and / or providing an input to discard an object. In some embodiments, while the device detects that the user's line of sight is directed at a selectable object (e.g., a photo file icon, a video file icon, a notification banner, etc.) within a three-dimensional environment, the device detects a thumb touch-down on the user's index finger followed by an upward wrist flick gesture, and the device activates an experience corresponding to the object (e.g., opens a photo in the air, starts a 3D movie, opens an expanded notification, etc.).
[0219] In some embodiments, in response to detecting, using one or more cameras, movement of a user's thumb on the user's index finger (e.g., as opposed to more exaggerated gestures such as waving a finger or hand in the air or sliding on a touch-sensing surface), and in accordance with a determination that the movement includes a touch down of the thumb on the index finger of a first hand followed by a hand rotation gesture of the first hand (e.g., a rotation of at least a portion of the first hand relative to the wrist of the first hand), the computer system performs an eleventh operation that is different from a first operation (e.g., different from each or a subset of the first through tenth operations corresponding to other types of movement patterns of the user's finger). For example, the eleventh operation adjusts an amount, value corresponding to an amount of hand rotation (e.g., rotates a virtual object, or a user interface object (e.g., a virtual dial control) selected and / or held by the user (e.g., using eye gaze), in accordance with the hand rotation gesture).
[0220] In some embodiments, while displaying a view of a three-dimensional environment, the computer system detects movement of the palm of a user's first hand toward the user's face. In accordance with a determination that the detected movement of the palm of the user's first hand toward the user's face meets a call criterion, the computer system performs a twelfth operation that is different from a first operation (e.g., different from each or a subset of the first through eleventh operations corresponding to other types of movement patterns of the user's finger). For example, the twelfth operation includes displaying a user interface object associated with a virtual assistant and / or displaying an image (e.g., a virtual representation of the user, a camera view of the user, a magnified view of the three-dimensional environment, and / or a magnified view of an object (e.g., a virtual object and / or a real object within the three-dimensional environment)) at a position corresponding to the palm of the first hand. In some embodiments, the call criterion includes a criterion that is met in accordance with a determination that the distance between the user's palm and the user's face has decreased below a threshold distance. In some embodiments, the call criterion includes a criterion that is met in accordance with a determination that the fingers of the hand are extended.
[0221] The specific order described for the operations in FIG. 8 is merely an example and is not intended to indicate that the described order is the only order in which the operations can be performed. One of ordinary skill in the art will recognize various ways to reorder the operations described herein. Additionally, it should be noted that the details of the other processes described herein with respect to the other methods described herein (e.g., methods 9000, 10000, 11000, 12000, and 13000) are also applicable in a manner similar to the method 8000 described above in connection with FIG. 8. For example, the contact, gesture, gaze input, physical object, and user interface object, and animation described above with reference to method 8000 optionally have one or more of the characteristics of the contact, gesture, gaze input, physical object, and user interface object, and animation described herein with reference to the other methods described herein (e.g., methods 9000, 10000, 11000, 12000, and 13000). For the sake of brevity, those details are not repeated here.
[0222] FIG. 9 is a flowchart of an exemplary method 9000 for interacting with a three-dimensional environment using a predetermined input gesture according to some embodiments. In some embodiments, method 9000 is executed in a computer system (e.g., computer system 101 of FIG. 1) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera that faces downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras)) or a camera that faces forward from the user's head). In some embodiments, method 9000 is stored on a non-transitory computer-readable storage medium and is executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some operations of method 9000 are optionally combined and / or the order of some operations is optionally changed.
[0223] In method 9000, a computer system displays (9002) a view of a three-dimensional environment (e.g., a virtual environment or an augmented reality environment). While displaying the three-dimensional environment, the computer system detects a hand at a first position corresponding to a portion of the three-dimensional environment (e.g., at a position in the physical environment where the hand is made visible to the user according to the current field of view of the user in the three-dimensional environment (e.g., the user's hand is moving to a position that intersects or is near the user's line of sight)). In some embodiments, a representation or image of the user's hand is displayed as part of the three-dimensional environment in response to detecting the hand at the first position within the physical environment. In response to detecting a hand at a first position corresponding to a portion of the three-dimensional environment (9004) and in accordance with a determination that the hand is maintained in a first predetermined configuration (e.g., using a camera or a touch-sensing glove or a touch-sensing finger attachment to detect a predetermined ready state such as the thumb being placed on the index finger), a visual indication of a first operating context for gesture input using hand gestures within the three-dimensional environment (e.g., system affordances (e.g., system affordance 7214 in FIGS. 7E, 7F, and 7G), docks, menus, visual indications such as avatars for voice-based virtual assistants, display of additional information regarding user interface elements within the three-dimensional environment that can be operated in response to hand gesture input, etc.) is displayed (e.g., proximate to the representation of the hand that is part of the three-dimensional environment), and in accordance with a determination that the hand is not maintained in the first predetermined configuration, the display of the visual indication of the first operating context for gesture input using hand gestures within the three-dimensional environment is not performed (e.g., as shown in FIG. 7D, the representation of the hand is displayed as part of the three-dimensional environment without displaying the visual indication proximate to the representation of the hand).
[0224] In some embodiments, a visual indication of a first operating context for gesture input using hand gestures is displayed at a location of a portion of the three-dimensional environment corresponding to a first location (e.g., the detected hand location). For example, the visual indication (e.g., a home affordance or a dock, etc.) is displayed at a location at and / or within a predetermined distance of the detected hand location. In some embodiments, the visual indication is displayed at a location corresponding to a particular portion of the hand (e.g., above the upper part of the detected hand, below the lower part of the detected hand, and / or overlaid on the hand).
[0225] In some embodiments, while displaying a visual indication in a portion of the three-dimensional environment, the computer system detects a change in the hand location from a first location to a second location (e.g., detects movement and / or rotation of the hand within the three-dimensional environment (e.g., while the hand is in a first predetermined configuration, or some other predetermined configuration indicating readiness of the hand)). In response to detecting a change in the hand location from the first location to the second location, the computer system changes the displayed location of the visual indication according to the detected change in the hand location (e.g., maintains the display of the visual indication within a predetermined distance of the hand within the three-dimensional environment).
[0226] In some embodiments, the visual indication includes a set of one or more user interface objects. In some embodiments, the visual indicator is a system affordance icon (e.g., system affordance 7120 in FIG. 7E) that indicates a region where one or more user interface objects can be displayed and / or accessed. For example, as shown in part (A) of FIG. 7F, when the thumb of hand 7200 moves across the index finger in direction 7120, the display of visual indicator 7214 is replaced by the display of a set of user interface objects 7170. In some embodiments, the presence of a system affordance icon near the user's hand indicates that the next gesture provided by the hand will trigger a system-level operation (e.g., an operation that is not object- or application-specific, an operation executed by the operating system independently of the application), and the absence of a system affordance icon near the user's hand indicates that the next gesture provided by the hand will trigger an application- or object-specific operation (e.g., an operation that is specific to or occurs within the currently selected object or application, e.g., an operation executed by the application). In some embodiments, the device displays a system affordance near the hand according to a determination that the user's line of sight is directed at the hand in a ready state.
[0227] In some embodiments, the one or more user interface objects include a plurality of application launch icons (e.g., the one or more user interface objects is a dock that includes a row of application launch icons for a plurality of frequently used applications or experiences), and activation of each of the application launch icons causes an operation associated with the corresponding application to be executed (e.g., launches the corresponding application).
[0228] In some embodiments, while displaying the visual indication, the computer system detects a change in the hand configuration from a first predetermined configuration to a second predetermined configuration (e.g., detecting a change in the position of the thumb relative to another finger such as a movement across another finger). In response to detecting the detected change in the hand configuration from the first predetermined configuration to the second predetermined configuration, the computer system displays a first set of user interface objects (e.g., a home region or an application launch user interface) (e.g., in addition to and / or in exchange for the visual indication), and activation of each user interface object of the first set of user interface objects causes an operation associated with each user interface object to be performed. In some embodiments, the visual indicator is a system affordance icon (e.g., system affordance 7214 of FIGS. 7E and 7F) that indicates a region where the home region or the application launch user interface can be displayed and / or accessed. For example, as shown in part (A) of FIG. 7F, when the thumb of hand 7200 moves in direction 7120 across the index finger, the display of visual indicator 7214 is replaced with the display of a set of user interface objects 7170. In some embodiments, at least some of the user interface objects of the first set of user interface objects are application launch icons, and activation of the application launch icons launches the corresponding applications.
[0229] In some embodiments, while displaying the visual indication, the computer system determines whether hand movement satisfies an interaction criterion (e.g., the interaction criterion is satisfied according to a determination that at least one finger and / or thumb of the hand moves a distance exceeding a threshold distance and / or moves according to a predetermined gesture) during a time window (e.g., a time window of 5 seconds, 8 seconds, 15 seconds, etc. from the time the visual indication is displayed in response to detecting a ready hand at a first position). According to a determination that the hand movement does not satisfy the interaction criterion during the time window, the computer system stops displaying the visual indication. In some embodiments, the device redisplayed the visual indication when the user's hand is detected again in a first predetermined configuration within the user's field of view after the user's hand has left the user's field of view or the user's hand has changed to another configuration that is not the first or other predetermined configuration corresponding to the ready state of the hand.
[0230] In some embodiments, while displaying the visual indication, the computer system detects a change in hand configuration from a first predetermined configuration to a second predetermined configuration that meets the input criteria (e.g., the hand configuration has changed, but the hand remains within the user's field of view). For example, the detected change can be a change in the position of the thumb (e.g., a change relative to another finger such as contact with another finger and / or release of contact from another finger, movement along the length of another finger, and / or movement across another finger), and / or a change in the position of the index finger and / or middle finger of the hand (e.g., extension of the finger and / or other movement of the finger relative to the hand). In response to detecting a change in hand configuration from a first predetermined configuration to a second configuration that meets the input criteria (e.g., according to a determination that the user's hand has changed from a configuration that is the starting state of a first accepted gesture to a configuration that is the starting state of a second accepted gesture), the computer system adjusts the visual indication (e.g., adjusts each selected user interface object of a set of one or more user interface objects from a first respective user interface object to a second respective user interface object, changes the displayed position of one or more user interface objects, and / or displays and / or stops displaying each respective user interface object of one or more user interface objects).
[0231] In some embodiments, while displaying the visual indication, the computer system detects a change in the hand configuration from a first predetermined configuration to a third configuration that does not meet the input criteria (e.g., according to a determination that at least a portion of the hand is outside the user's field of view, the configuration does not meet the input criteria). In some embodiments, the device determines that the third configuration does not meet the input criteria according to a determination that the user's hand has changed from a configuration that is the starting state of a first accepted gesture to a state that does not correspond to the starting state of any accepted gesture. In response to detecting a detected change in the hand configuration from the first predetermined configuration to a third configuration that does not meet the input criteria, the computer system stops displaying the visual indication.
[0232] In some embodiments, after stopping the display of the visual indication, the computer system detects a change in the hand configuration (and that the hand is within the user's field of view) to the first predetermined configuration. In response to detecting the detected change in the hand configuration to the first predetermined configuration, the computer system redisplay the visual indication.
[0233] In some embodiments, in response to detecting a hand at a first position corresponding to a portion of a three-dimensional environment, according to a determination that the hand is not maintained in the first predetermined configuration, the computer system performs an operation different from displaying a visual indication of a first operating context for gesture input using the hand gesture (e.g., displaying a representation of the hand without a visual indication and / or providing a prompt indicating that the hand is not maintained in the first predetermined configuration).
[0234] The specific order described for the operations in FIG. 9 is merely an example and is not intended to indicate that the described order is the only order in which the operations can be performed. One of ordinary skill in the art will recognize various ways to reorder the operations described herein. Additionally, it should be noted that the details of the other processes described herein with respect to the other methods described herein (e.g., methods 8000, 10000, 11000, 12000, and 13000) are also applicable in a manner similar to the method 9000 described above in connection with FIG. 9. For example, the contact, gesture, eye gaze input, physical object, and user interface object, and animation described above with reference to method 9000 optionally have one or more of the characteristics of the contact, gesture, eye gaze input, physical object, and user interface object, and animation described herein with reference to the other methods described herein (e.g., methods 8000, 10000, 11000, 12000, and 13000). For the sake of brevity, those details are not repeated here.
[0235] FIG. 10 is a flowchart of an exemplary method 10000 for interacting with a three-dimensional environment using a predetermined input gesture, according to some embodiments. In some embodiments, method 10000 is executed in a computer system (e.g., computer system 101 of FIG. 1) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera facing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth sensing cameras)) or a camera facing forward from the user's head). In some embodiments, method 10000 is stored in a non-transitory computer-readable storage medium and is executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some operations of method 10000 are optionally combined and / or the order of some operations is optionally changed.
[0236] In method 10000, a computer system displays a three-dimensional environment (e.g., an augmented reality environment) that includes displaying a representation of the physical environment (e.g., displaying a camera view of the physical environment surrounding the user, or including a pass-through portion that displays the physical environment surrounding the user within the displayed user interface or virtual environment) (10002). While displaying the representation of the physical environment, the computer system detects gestures (e.g., gestures including a predetermined movement of the user's hand, finger, wrist, or arm, or a predetermined static pose of the hand that is different from the natural resting pose of the hand) (e.g., using a camera or one or more motion sensors) (10004). In response to detecting a gesture (10006), and according to a determination that the user's line of sight is directed to a position (e.g., within the three-dimensional environment) corresponding to a predetermined physical position (e.g., the user's hand) within the physical environment (e.g., according to a determination that the line of sight is directed to and remains at that position during the start and completion of the gesture, or according to a determination that the line of sight is directed to the hand while the hand is in the final state of the gesture (e.g., the ready state of the hand (e.g., a predetermined static pose of the hand))), the computer system displays a system user interface within the three-dimensional environment (e.g., a user interface including visual indications of interaction options available in the three-dimensional environment and / or selectable options, the user interface being displayed in response to the gesture and not being displayed before detection of the gesture (e.g., when the line of sight was directed)). This is shown, for example, in parts A-1, A-2, and A-3 of FIG. 7G where an input gesture by hand 7200 causes interaction with system user interface elements such as system affordance 7214, system menu 7170, and application icon 7190. In some embodiments, the position corresponding to the predetermined physical position is a representation (e.g., a video image or graphical abstraction) of a predetermined physical position within the three-dimensional environment (the user's hand that is movable within the physical environment, or a physical object that is stationary within the physical environment).In some embodiments, the system user interface includes one or more application icons (e.g., when activated, each application icon launches the corresponding application). In response to detecting a gesture (10006), and in accordance with a determination that the user's line of sight is not directed towards a position corresponding to a predetermined physical position within the physical environment (e.g., within a three-dimensional environment) (e.g., in accordance with a determination that the line of sight is directed towards another position and / or remains at another position, or that no line of sight is detected during the start and completion of the gesture, or in accordance with a determination that the line of sight is not directed towards the hand while the hand is in the final state of the gesture (e.g., the ready state of the hand (e.g., a predetermined stationary posture of the hand))), an operation within the current context of the three-dimensional environment is performed without displaying the system user interface. This is shown, for example, in FIGS. 7G, B-1, B-2, and B-3. In some embodiments, the operation includes a first operation (e.g., changing the output volume of the device) that changes the state of the electronic device without generating a visual change in the three-dimensional environment. In some embodiments, the operation includes a second operation that displays the hand making the gesture and does not cause further interaction with the three-dimensional environment. In some embodiments, the operation includes an operation for changing the state of a virtual object based on the fact that the line of sight is currently directed towards it. In some embodiments, the operation includes an operation for changing the state of the virtual object with which the user last interacted in the three-dimensional environment. In some embodiments, the operation includes an operation for changing the state of the currently selected virtual object and having an input focus.
[0237] In some embodiments, the computer system displays a home affordance (e.g., a system affordance for system-level operations (as opposed to the application level)) that indicates that it is ready to detect one or more system gestures for displaying a user interface at a location corresponding to a given physical location. In some embodiments, the location corresponding to the given physical location is a location within a three-dimensional environment. In some embodiments, the location corresponding to the given physical location is a location on a display. In some embodiments, as long as the location within the three-dimensional environment where the given location of the system affordance is displayed, the system affordance remains displayed even if the location corresponding to the given physical location is no longer visible in the displayed portion of the three-dimensional environment (e.g., the system affordance continues to be displayed even if the given physical location goes out of the field of view of one or more cameras of the electronic device). In some embodiments, the system affordance is displayed at a given location relative to a given fixed position of the user's hand, wrist, or finger with respect to a representation of the user's hand, wrist, or finger within the three-dimensional environment (e.g., superimposed or replaced on a display of a part of the user's hand, wrist, or finger or at a fixed position offset from the user's hand, wrist, or finger). In some embodiments, the system affordance is displayed at a given location relative to the location corresponding to the given physical location regardless of whether the user's line of sight remains directed at the location within the three-dimensional environment (e.g., the system affordance remains displayed within a given timeout period even after the user's line of sight has moved away from the user's hand in a ready state or after a gesture has been completed).
[0238] In some embodiments, displaying a system affordance at a location corresponding to a given physical location with respect to a location means detecting movement at a location corresponding to the given physical location within the three-dimensional environment (e.g., detecting that the position of the user's hand shown in the three-dimensional environment has changed as a movement of the user's head or hand), and in response to detecting movement at a location corresponding to the given physical location within the three-dimensional environment, moving the system affordance within the three-dimensional environment so that the relative position of the system affordance and the location corresponding to the given physical location remain unchanged in the three-dimensional environment (e.g., when the position of the user's hand changes within the three-dimensional environment, the system affordance follows the position of the user's hand (e.g., the system affordance is displayed at a position corresponding to above the user's thumb in the displayed view of the three-dimensional environment)).
[0239] In some embodiments, the system affordance is displayed at a given position with respect to a position corresponding to a given physical location according to a determination that the user's line of sight is directed at the position corresponding to the given physical location. In some embodiments, the system affordance is displayed at a given relative position according to a determination that the user's line of sight is directed at a position in the vicinity of the given physical location (e.g., within a given threshold distance of the given physical location). In some embodiments, when the user's line of sight is not directed at the given physical location (e.g., when the user's line of sight is away from the given physical location or at least a given distance away from the given physical location), the system affordance is not displayed. In some embodiments, while the system affordance is being displayed at a given position with respect to a position corresponding to a given physical location within the three-dimensional environment, the device detects that the user's line of sight has moved away from the position corresponding to the given physical location, and in response to detecting that the user's line of sight has moved away from the position corresponding to the given physical location within the three-dimensional environment, the device stops displaying the system affordance at the given position within the three-dimensional environment.
[0240] In some embodiments, displaying a system affordance at a location corresponding to a predetermined physical location within a three-dimensional environment includes displaying a system affordance having a first appearance (e.g., shape, size, color, etc.) in accordance with a determination that the user's line of sight is not directed at the location corresponding to the predetermined physical location, and displaying a system affordance having a second appearance different from the first appearance in accordance with a determination that the user's line of sight is directed at the location corresponding to the predetermined physical location. In some embodiments, the system affordance has the first appearance while the user's line of sight is away from the location corresponding to the predetermined physical location. In some embodiments, the system affordance has the second appearance while the user's line of sight is directed at the location corresponding to the predetermined physical location. In some embodiments, the system affordance changes from the first appearance to the second appearance when the user's line of sight moves towards (e.g., within at least a threshold distance of) the location corresponding to the predetermined physical location, and changes from the second appearance to the first appearance when the user's line of sight moves away from (e.g., at least a threshold distance away from) the location corresponding to the predetermined physical location.
[0241] In some embodiments, the system affordance is displayed at a predetermined position relative to a position corresponding to a predetermined physical position, in accordance with a determination that the user is ready to perform a gesture. In some embodiments, determining that the user is ready to perform a gesture includes detecting an indication that the user is ready to perform a gesture, such as by detecting that a predetermined physical position (e.g., the user's hand, wrist, or finger(s)) is in a predetermined configuration (e.g., a predetermined orientation relative to a device in the physical environment). In one example, the system affordance is displayed at a predetermined position relative to a representation of the user's hand displayed in a three-dimensional environment when the device detects that the user has placed their hand in a predetermined ready state (e.g., a particular position and / or orientation of the hand) in the physical environment, in addition to detecting a line of sight to the ready hand. In some embodiments, the predetermined configuration requires that a predetermined physical position (e.g., the user's hand) has a particular position relative to an electronic device or one or more input devices of the electronic device, such as being within the field of view of one or more cameras.
[0242] In some embodiments, displaying a system affordance at a location corresponding to a given physical location includes displaying a system affordance having a first appearance according to a determination that the user is not ready to perform a gesture. In some embodiments, determining that the user is not ready to perform a gesture includes detecting an indication that the user is not ready to perform a gesture (e.g., detecting that the user's hand is not in a given ready state). In some embodiments, determining that the user is not ready includes failing to detect an indication that the user is ready (e.g., if the user's hand is outside the field of view of one or more cameras of the electronic device, failing to detect or being unable to detect that the user's hand is in a given ready state). Detecting an indication of the user's readiness to perform a gesture is described in more detail herein with reference to FIG. 7E and the related description. In some embodiments, displaying a system affordance at a location corresponding to a given physical location further includes displaying a system affordance having a second appearance different from the first appearance according to a determination that the user is ready to perform a gesture (e.g., according to detecting an indication that the user is ready to perform a gesture as described herein with reference to FIG. 7E and the related description). One of ordinary skill in the art will recognize that the presence or absence of a system affordance, and the particular appearance of the system affordance, may be modified according to what information is intended to be conveyed to the user in that particular context (e.g., what action(s) will be performed in response to the gesture and / or whether additional criteria need to be met for the gesture to invoke the system user interface).In some embodiments, while a system affordance is displayed at a given position corresponding to a given physical position within a three-dimensional environment, the device detects that the user's hand changes from a first state to a second state. In response to detecting the change from the first state to the second state, according to the determination that the first state is a ready state and the second state is not a ready state, the device displays a system affordance having a second appearance (a change from the first appearance). According to the determination that the first state is not a ready state and the second state is a ready state, the device displays a system affordance having a first appearance (e.g., a change from the second appearance). In some embodiments, when a computer system does not detect the user's line of sight on the user's hand and the user's hand is not in a ready state configuration, the computer system does not display the system affordance, or optionally, displays a system affordance having a first appearance. When a subsequent input gesture is detected (e.g., when the system affordance is not displayed or is displayed with a first appearance), the computer system does not perform a system operation corresponding to the input gesture, or optionally, performs an operation within the current user interface context corresponding to the input gesture. In some embodiments, when a computer system detects the user's line of sight on the user's hand and the hand is not in a ready state configuration, the computer system does not display the system affordance, or optionally, displays a system affordance having a first appearance or a second appearance. When a subsequent gesture input is detected (e.g., when the system affordance is not displayed or is displayed with a first appearance or a second appearance), the computer system does not perform a system operation corresponding to the input gesture, or optionally, displays a system user interface (e.g., a dock or a system menu).In some embodiments, the computer system does not detect the user's line of sight on the user's hand, but if the hand is in a ready state configuration, the computer system does not display system affordances, or optionally displays system affordances having a first appearance or a second appearance. If subsequent gesture input is detected (e.g., when the system affordances are not displayed, or are displayed with the first appearance or the second appearance), the computer system does not perform system operations corresponding to the input gesture, or optionally performs operations within the current user interface context. In some embodiments, the computer system detects the user's line of sight on the user's hand, and if the hand is in a ready state configuration, the computer system displays system affordances having a second appearance or a third appearance. If subsequent gesture input is detected (e.g., when the system affordances have the second appearance or the third appearance), the computer system performs system operations (e.g., displays the system user interface). In some embodiments, a plurality of the above are combined in the same embodiment.
[0243] In some embodiments, the predetermined physical location is the user's hand, and determining that the user is ready to perform a gesture (e.g., the hand is currently in a predetermined ready state or a start gesture has been detected) includes determining that a predetermined portion of the hand (e.g., a specified finger) is in contact with a physical control element. In some embodiments, the physical control element is a controller separate from the user (e.g., respective input devices) (e.g., the ready state is a state in which the user's thumb is in contact with a touch-sensing strip or ring attached to the user's index finger). In some embodiments, the physical control element is a different portion of the user's hand (e.g., the ready state is a state in which the thumb is in contact with the upper side of the index finger (e.g., near the second joint)). In some embodiments, the device uses a camera to detect whether the hand is in a predetermined ready state and displays the hand in the ready state in a view of the three-dimensional environment. In some embodiments, the device is touch-sensitive and uses a physical control element communicatively coupled to the electronic device to transmit touch input to the electronic device to detect whether the hand is in a predetermined ready state.
[0244] In some embodiments, the predetermined physical location is the user's hand, and determining that the user is ready to perform a gesture includes determining that the hand is raised above a predetermined altitude relative to the user. In some embodiments, determining that the hand is raised includes determining that the hand is positioned above a particular lateral plane relative to the user (e.g., above the user's waist, i.e., closer to the user's head than the user's feet). In some embodiments, determining that the hand is raised includes determining that the user's wrist or elbow is bent at least a particular amount (e.g., within an angle of 90 degrees). In some embodiments, the device uses a camera to detect whether the hand is in a predefined ready state and optionally displays the hand in the ready state in a view of the three-dimensional environment. In some embodiments, the device uses one or more sensors (e.g., motion sensors) attached to the user's hand, wrist, or arm to detect whether the hand is in a predetermined ready state and is communicatively coupled to an electronic device to transmit a movement input to the electronic device.
[0245] In some embodiments, the predetermined physical location is the user's hand, and determining that the user is ready to perform a gesture includes determining that the hand is in a predetermined configuration. In some embodiments, the predetermined configuration requires that a corresponding finger of the hand (e.g., the thumb) is in contact with a different part of the user's hand (e.g., an opposing finger such as the index finger, or a predetermined part such as the middle phalanx or middle joint of the opposing index finger). In some embodiments, the predetermined configuration requires that the hand is above a particular lateral plane as described above (e.g., above the user's waist). In some embodiments, the predetermined configuration requires flexion of the wrist away from the little finger side towards the thumb side (e.g., radial flexion) (e.g., without axial rotation of the arm). In some embodiments, when the hand is in the predetermined configuration, one or more fingers are in a natural rest position (e.g., curved), and the hand as a whole is tilted or moved away from the natural rest position of the hand, wrist, or arm, indicating that the user is ready to perform a gesture. One of ordinary skill in the art will recognize that the particular predetermined ready state used can be selected to have an intuitive and natural user interaction and may require any combination of the foregoing criteria. In some embodiments, when the user simply desires to view a three-dimensional environment rather than provide input to and interact with the three-dimensional environment, the predetermined configuration is different from the natural rest posture of the user's hand (e.g., a posture of relaxation placed on the knee, table surface, or side of the body). The change from the natural rest posture to the predetermined configuration is intentional and requires an intentional movement of the user's hand to the predetermined configuration.
[0246] In some embodiments, the position corresponding to a given physical location is a fixed position within the three-dimensional environment (e.g., the corresponding given physical location is a fixed position within the physical environment). In some embodiments, the physical environment is the user's reference frame. That is, one of ordinary skill in the art will recognize that a position referred to as a "fixed" position within the physical environment may not be an absolute position in space, but rather is fixed with respect to the user's frame of reference. In some examples, when the user is inside a building room, the position is a fixed position within the three-dimensional environment corresponding to (e.g., represented by) a fixed position within the room (e.g., on the walls, floor, or ceiling of the room). In some examples, when the user is inside a moving vehicle, the position is a fixed position within the three-dimensional environment corresponding to (e.g., representing) a fixed position along the interior of the vehicle. In some embodiments, the position is fixed with respect to the content displayed in the three-dimensional environment, and the displayed content corresponds to a fixed given physical location within the physical environment.
[0247] In some embodiments, the position corresponding to a given physical location is a fixed position with respect to the display of the three-dimensional environment (e.g., with respect to the display generation component). In some embodiments, the position is fixed with respect to the user's viewpoint of the three-dimensional environment regardless of the particular content displayed within the three-dimensional environment that is generally updated as the user's viewpoint changes (e.g., in response to or along with the user's viewpoint change) (e.g., the position is fixed with respect to the display of the three-dimensional environment by the display generation component). In some examples, the position is a fixed position along the edge of the display of the three-dimensional environment (e.g., within a given distance of the edge). In some examples, the position is centered with respect to the display of the three-dimensional environment (e.g., centered within a display area along the bottom, top, left edge, or right edge of the display of the three-dimensional environment).
[0248] In some embodiments, the predetermined physical location is a fixed location on the user. In some examples, the predetermined physical location is the user's hand or finger. In some such examples, the location corresponding to the predetermined physical location includes a displayed representation of the user's hand or finger within the three-dimensional environment.
[0249] In some embodiments, after displaying the system user interface in a three-dimensional environment, the computer system detects a second gesture (e.g., a second gesture performed by the user's hand, wrist, finger(s), or arm) while (e.g., after detecting a line of sight directed to a position corresponding to a first gesture and a predetermined physical position and while the system user interface is being displayed). In response to detecting the second gesture, the system user interface (e.g., an application launch user interface) is displayed. In some embodiments, the second gesture is a continuation of the first gesture. For example, the first gesture is a swipe gesture (e.g., by moving the user's thumb on the user's index finger with the same hand), and the second gesture is a continuation of the swipe gesture (e.g., a continuous movement of the thumb on the index finger) (e.g., the second gesture starts from the end position of the first gesture without resetting the start position of the second gesture to the start position of the first gesture). In some embodiments, the second gesture is a repetition of the first gesture (e.g., after performing the first gesture, the start position of the second gesture is reset within a predetermined distance of the start position of the first gesture, and the second gesture retraces the movement of the first gesture within a predetermined tolerance). In some embodiments, displaying the home user interface includes expanding system affordances from a predetermined position with respect to a position corresponding to a predetermined physical position to occupy a larger portion of the displayed three-dimensional environment and showing additional user interface objects and options.In some embodiments, the system affordance is an indicator that does not include respective content, and each content (e.g., a dock having an application icon list of recently used or frequently used applications) replaces the indicator in response to a first swipe gesture by hand, a two-dimensional grid of all application icons of installed applications replaces the dock in response to a second swipe gesture by hand, and a three-dimensional work environment having interactive application icons floating at different depths and positions within the three-dimensional work environment replaces the two-dimensional grid in response to a third swipe gesture by hand.
[0250] In some embodiments, the current context of the three-dimensional environment includes the display of an indication of a received notification (e.g., an initial display of a subset of information regarding the received notification), and performing an operation within the current context of the three-dimensional environment includes displaying an expanded notification that includes additional information regarding the received notification (e.g., displaying information beyond the initially displayed subset). In some embodiments, the current context of the three-dimensional environment is determined based on the position at which the line of sight is currently directed. In some embodiments, when a notific...
Claims
**Claim 1** A method comprising: in a computer system including a display generation component and one or more cameras, displaying a view of a three-dimensional environment; while displaying the view of the three-dimensional environment, using the one or more cameras to detect movement of a user's thumb on the user's index finger of a first hand of the user; in response to detecting, using the one or more cameras, the movement of the user's thumb on the user's index finger, performing a first action according to a determination that the movement is a swipe of the thumb on the index finger of the first hand in a first direction; performing a second action different from the first action according to a determination that the movement is a tap of the thumb on the index finger at a first position on the index finger of the first hand; A method comprising the above. **Claim 2** In response to detecting, using the one or more cameras, the movement of the user's thumb on the user's index finger, performing a third action different from the first action and the second action according to a determination that the movement is a swipe of the thumb on the index finger of the first hand in a second direction substantially perpendicular to the first direction. The method according to claim 1. **Claim 3** In response to detecting, using the one or more cameras, the movement of the user's thumb on the user's index finger, performing a fourth action different from the first action and the second action according to a determination that the movement is a movement of the thumb on the index finger in a third direction different from the first direction. The method according to claim 1 or 2. **Claim 4** Performing the first action includes: increasing a value corresponding to the first action according to a determination that the swipe of the thumb on the index finger in the first direction moves towards a first predetermined portion of the index finger; decreasing the value corresponding to the first action according to a determination that the swipe of the thumb on the index finger in the first direction moves away from the first predetermined portion of the index finger and towards a second predetermined portion of the index finger; The method according to any one of claims 1 to 3. **Claim 5** Performing the first operation includes adjusting a value corresponding to the first operation by an amount corresponding to an amount of movement of the thumb on the index finger, according to any one of claims 1 to 4.
6. In response to detecting, using the one or more cameras, the movement of the user's thumb on the user's index finger, performing a fifth operation different from the second operation according to a determination that the movement is a tap of the user's thumb on the index finger at a second position different from the first position on the index finger, according to any one of claims 1 to 5.
7. Detecting, using the one or more cameras, a swipe of the user's thumb on the user's middle finger, and in response to detecting the swipe of the user's thumb on the user's middle finger, performing a sixth operation, according to any one of claims 1 to 6.
8. Detecting, using the one or more cameras, a tap of the user's thumb on the user's middle finger, and in response to detecting the tap of the user's thumb on the user's middle finger, performing a seventh operation, according to any one of claims 1 to 7.
9. Including displaying a visual indication of an operation context of a thumb gesture in the three-dimensional environment, and performing the first operation includes displaying a visual change in the three-dimensional environment corresponding to the execution of the first operation, according to any one of claims 1 to 8.
10. While displaying the visual indication of the operation context of the thumb gesture, detecting, using the one or more cameras, movement of the user's first hand, and in response to detecting the movement of the first hand, changing a displayed position of the visual indication of the operation context of the thumb gesture in the three-dimensional environment according to the detected change in the position of the hand, according to claim 9.
11. Detecting, using the one or more cameras, movement of the user's thumb on the user's index finger of the user's second hand, and in response to detecting, using the one or more cameras, the movement of the user's thumb on the user's index finger of the second hand, Execute an eighth operation different from the first operation according to the determination that the movement is a swipe of the thumb on the index finger of the second hand in the first direction. Execute a ninth operation different from the second operation according to the determination that the movement is a tap of the thumb on the index finger of the second hand at the first position on the index finger of the second hand. The method according to any one of claims 1 to 10, including the above.
12. In response to detecting the movement of the user's thumb on the user's index finger using the one or more cameras, execute a tenth operation different from the first operation according to the determination that the movement includes a touch - down of the thumb on the index finger of the first hand and subsequently a wrist - flick gesture of the first hand. The method according to any one of claims 1 to 11, including the above.
13. In response to detecting the movement of the user's thumb on the user's index finger using the one or more cameras, execute an eleventh operation different from the first operation according to the determination that the movement includes a touch - down of the thumb on the index finger of the first hand and subsequently a hand - rotation gesture of the first hand. The method according to any one of claims 1 to 11, including the above.
14. While displaying the view of the three - dimensional environment, detect the movement of the palm of the user's first hand towards the user's face. Execute a twelfth operation different from the first operation according to the determination that the movement of the palm of the user's first hand towards the user's face meets a call criterion. The method according to any one of claims 1 to 13, including the above. The method according to any one of claims 1 to 13, including the above.
15. A computer - readable storage medium storing executable instructions that, when executed by a computer system having one or more processors and a display - generation component, cause the computer system to execute the method according to any one of claims 1 to 14.
16. A computer system, One or more processors, A display - generation component, A memory storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs include instructions for executing the method according to any one of claims 1 to 14, a memory, A computer system comprising.
17. A computer system comprising one or more processors and a display generation component, A computer system comprising means for executing the method according to any one of claims 1 to 14.
18. An information processing apparatus for use in a computer system comprising one or more processors and a display generation component, An information processing apparatus comprising means for executing the method according to any one of claims 1 to 14.
19. A method, In a computing system including a display generation component and one or more input devices, Displaying a view of a three-dimensional environment; While displaying the three-dimensional environment, detecting a hand at a first position corresponding to a part of the three-dimensional environment; In response to detecting the hand at the first position corresponding to the part of the three-dimensional environment, According to the determination that the hand is maintained in a first predetermined configuration, displaying a visual indication of a first operation context for gesture input using hand gestures within the three-dimensional environment; According to the determination that the hand is not maintained in the first predetermined configuration, not displaying the visual indication of the first operation context for gesture input using hand gestures within the three-dimensional environment; A method comprising.
20. The method according to claim 19, wherein the visual indication of the first operation context for gesture input using hand gestures is displayed at a position of the part of the three-dimensional environment corresponding to the first position.
21. While displaying the visual indication on the part of the three-dimensional environment, detecting a change in the position of the hand from the first position to a second position; In response to detecting the change in the position of the hand from the first position to the second position, changing the displayed position of the visual indication according to the detected change in the position of the hand; The method according to claim 20, comprising.
22. The visual indication includes one or more user interface objects, and the method according to any one of claims 19 to 21.
23. The method according to claim 22, wherein the one or more user interface objects include a plurality of application launch icons, and activation of each of the application launch icons causes an operation associated with the corresponding application to be executed.
24. While displaying the visual indication, detecting a change in the configuration of the hand from the first predetermined configuration to the second predetermined configuration; In response to detecting the change in the configuration of the hand from the first predetermined configuration to the second predetermined configuration, displaying a first set of user interface objects, wherein activation of each user interface object in the first set of user interface objects causes an operation associated with the respective user interface object to be executed. The method according to any one of claims 19 to 21, comprising:
25. While displaying the visual indication, determining whether the movement of the hand satisfies an interaction criterion within a time window; Stopping the display of the visual indication according to the determination that the movement of the hand does not satisfy the interaction criterion within the time window. The method according to claim 24, comprising:
26. While displaying the visual indication, detecting a change in the hand configuration from the first predetermined configuration to a second predetermined configuration that satisfies an input criterion; Adjusting the visual indication in response to detecting the change in the hand configuration from the first predetermined configuration to the second configuration that satisfies the input criterion. The method according to any one of claims 19 to 25, comprising:
27. While displaying the visual indication, detecting a change in the hand configuration from the first predetermined configuration to a third configuration that does not satisfy the input criterion; Stopping the display of the visual indication in response to detecting the change in the configuration of the hand from the first predetermined configuration to the third configuration that does not satisfy the input criterion. The method according to claim 26, comprising **Claim 28** After stopping displaying the visual indication, detecting a change in the configuration of the hand to the first predetermined configuration; In response to detecting the change in the configuration of the hand to the first predetermined configuration, redisplaying the visual indication; The method according to claim 27, comprising **Claim 29** In response to detecting the hand at the first position corresponding to the part of the three-dimensional environment, performing an operation different from displaying the visual indication of the first operation context for gesture input using a hand gesture according to a determination that the hand is not maintained in the first predetermined configuration. The method according to any one of claims 19 to 28. **Claim 30** A computer-readable storage medium storing executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to execute the method according to any one of claims 19 to 29. **Claim 31** A computer system, comprising One or more processors; A display generation component; A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for executing the method according to any one of claims 19 to 29; A computer system comprising **Claim 32** A computer system comprising one or more processors and a display generation component, The computer system comprising means for executing the method according to any one of claims 19 to 29. **Claim 33** An information processing apparatus for use in a computer system comprising one or more processors and a display generation component, The information processing apparatus comprising means for executing the method according to any one of claims 19 to 29. **Claim 34** A method, comprising In a computer system including a display generation component and one or more input devices, Displaying a three-dimensional environment including displaying a representation of a physical environment; Detecting a gesture while displaying the representation of the physical environment; In response to detecting the gesture, Displaying a system user interface in the three-dimensional environment according to a determination that the user's line of sight is directed toward a position corresponding to a predetermined physical position in the physical environment; Executing an operation in the current context of the three-dimensional environment without displaying the system user interface according to a determination that the user's line of sight is not directed toward the position corresponding to the predetermined physical position in the physical environment; A method comprising the above.
35. The method according to claim 34, comprising displaying a system affordance at a predetermined position relative to the position corresponding to the predetermined physical position.
36. Displaying the system affordance at the predetermined position relative to the position corresponding to the predetermined physical position is Detecting a movement of the position corresponding to the predetermined physical position in the three-dimensional environment; In response to detecting a movement of the position corresponding to the predetermined physical position in the three-dimensional environment, moving the system affordance in the three-dimensional environment so that the relative position of the system affordance and the position corresponding to the predetermined physical position do not change in the three-dimensional environment; The method according to claim 35, comprising the above.
37. The method according to any one of claims 35 to 36, wherein the system affordance is displayed at the predetermined position relative to the position corresponding to the predetermined physical position according to a determination that the user's line of sight is directed toward the position corresponding to the predetermined physical position.
38. Displaying the system affordance at the predetermined position relative to the position corresponding to the predetermined physical position in the three-dimensional environment is Displaying the system affordance having a first appearance according to a determination that the user's line of sight is not directed toward the position corresponding to the predetermined physical position; Displaying the system affordance having a second appearance different from the first appearance according to a determination that the user's line of sight is directed toward the position corresponding to the predetermined physical position; The method according to claim 35 or 36, comprising the above.
39. The method according to claim 35 or 36, wherein the system affordance is displayed at the predetermined position with respect to the position corresponding to the predetermined physical position, in accordance with a determination that the user is ready to execute a gesture.
40. displaying the system affordance at the predetermined position with respect to the position corresponding to the predetermined physical position, displaying the system affordance having a first appearance, in accordance with a determination that the user is not ready to execute a gesture; displaying the system affordance having a second appearance different from the first appearance, in accordance with a determination that the user is ready to execute a gesture; The method according to claim 35 or 36, comprising:
41. The method according to claim 39 or 40, wherein the predetermined physical position is the user's hand, and determining that the user is ready to execute a gesture includes determining that a predetermined portion of the hand is in contact with a physical control element.
42. The method according to claim 39 or 40, wherein the predetermined physical position is the user's hand, and determining that the user is ready to execute a gesture includes determining that the hand is raised above a predetermined altitude relative to the user.
43. The method according to claim 39 or 40, wherein the predetermined physical position is the user's hand, and determining that the user is ready to execute a gesture includes determining that the hand is in a predetermined configuration.
44. The method according to any one of claims 34 to 43, wherein the position corresponding to the predetermined physical position is a fixed position within the three-dimensional environment.
45. The method according to any one of claims 34 to 43, wherein the position corresponding to the predetermined physical position is a fixed position with respect to the display of the three-dimensional environment.
46. The method according to any one of claims 34 to 43, wherein the predetermined physical position is a fixed position above the user.
47. after displaying the system user interface in the three-dimensional environment, detecting a second gesture; displaying the system user interface in response to detecting the second gesture; The method according to any one of claims 34 to 46, comprising:
48. The method according to any one of claims 34 to 47, wherein the current context of the three-dimensional environment includes displaying an indication of a received notification, and performing the operation in the current context of the three-dimensional environment includes displaying an expanded notification including additional information regarding the received notification.
49. The method according to any one of claims 34 to 47, wherein the current context of the three-dimensional environment includes displaying an indication of one or more photos, and performing the operation in the current context of the three-dimensional environment includes displaying at least one of the one or more photos within the three-dimensional environment.
50. A computer-readable storage medium storing executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to perform the method according to any one of claims 34 to 49.
51. A computer system, comprising one or more processors, a display generation component, a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 34 to 49, and a computer system.
52. A computer system comprising one or more processors and a display generation component, the computer system comprising means for performing the method according to any one of claims 34 to 49.
53. An information processing apparatus for use in a computer system comprising one or more processors and a display generation component, the information processing apparatus comprising means for performing the method according to any one of claims 34 to 49.
54. A method, in an electronic device including a display generation component and one or more input devices, displaying a three-dimensional environment including one or more virtual objects, detecting a line of sight directed at a first object within the three-dimensional environment, the line of sight satisfying a first criterion and the first object responding to at least one gesture input. In response to detecting the line of sight directed towards the first object that meets the first criterion and is responsive to at least one gesture input, displaying an indication of one or more interaction options available for the first object in the three-dimensional environment according to a determination that the hand is in a predetermined ready state for providing a gesture input; not displaying the indication of one or more interaction options available for the first object according to a determination that the hand is not in the predetermined ready state for providing a gesture input; A method comprising.
55. The method according to claim 54, wherein determining that the hand is in the predetermined ready state for providing a gesture input includes determining that a predetermined part of the hand is in contact with a physical control element.
56. The method according to claim 54, wherein determining that the hand is in the predetermined ready state for providing a gesture input includes determining that the hand is raised above the user to a predetermined height.
57. The method according to claim 54, wherein determining that the hand is in the predetermined ready state for providing a gesture input includes determining that the hand is in a predetermined configuration.
58. The method according to any one of claims 54 to 57, wherein displaying the indication of one or more interaction options available for the first object includes displaying information regarding the first virtual object that is adjustable in response to subsequent input.
59. The method according to any one of claims 54 to 58, wherein displaying the indication of one or more interaction options available for the first object includes displaying an animation of the first object.
60. The method according to any one of claims 54 to 59, wherein displaying the indication of one or more interaction options available for the first object includes displaying a selection indicator on at least a part of the first object.
61. The line of sight directed at a second object within the three-dimensional environment, wherein the line of sight meets the first criterion and the second virtual object responds to at least one gesture input, detecting the line of sight; In response to detecting the line of sight directed at the second virtual object that meets the first criterion and responds to at least one gesture input; Displaying an indication of one or more interaction options available for the second virtual object according to a determination that the hand is in the predetermined ready state for providing a gesture input; The method according to any one of claims 54 to 60, comprising:
62. In response to detecting the line of sight directed at the first object that meets the first criterion and responds to at least one gesture input; According to the determination that the hand is in the predetermined ready state for providing a gesture input; Detecting a first gesture input by the hand; Performing an interaction with the first object in response to detecting the first gesture input by the hand; The method according to any one of claims 54 to 61, comprising:
63. The method according to claim 62, wherein the first object includes a first image, and performing the interaction with the first object includes replacing the first image with a second image different from the first image.
64. The method according to claim 62, wherein the first object includes a first playable media content, and performing the interaction with the first object includes toggling the playback of the first playable media content.
65. The method according to claim 62, wherein the first object is a virtual window that displays a first virtual landscape, and performing the interaction with the first object includes replacing the display of the first virtual landscape with the display of a second virtual landscape different from the first virtual landscape.
66. The first gesture input is an upward flick gesture; The method according to claim 62, wherein performing the interaction with the first object includes displaying a user interface having one or more interaction options for the first object.
67. wherein the first gesture input includes rotation of the hand, and executing the interaction with the first object includes changing an output volume of content associated with the first object, The method according to claim 62.
68. The method according to any one of claims 54 to 67, wherein the first criterion includes a requirement that the line of sight remains directed at the first object for at least a threshold time.
69. A computer-readable storage medium storing executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to execute the method according to any one of claims 54 to 68.
70. A computer system, comprising one or more processors, a display generation component, a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for executing the method according to any one of claims 54 to 68, A computer system comprising.
71. A computer system comprising one or more processors and a display generation component, A computer system comprising means for executing the method according to any one of claims 54 to 68.
72. An information processing apparatus for use in a computer system comprising one or more processors and a display generation component, An information processing apparatus comprising means for executing the method according to any one of claims 54 to 68.
73. A method, in a computer system including a display generation component and one or more input devices, detecting an arrangement of the display generation component at a predetermined position with respect to a user of the computer system; in response to detecting the arrangement of the display generation component at the predetermined position with respect to the user of the computer system, displaying, through the display generation component, a first view of a three-dimensional environment including a pass-through portion, the pass-through portion including a representation of at least a part of the real world surrounding the user. While displaying the first view of the three-dimensional environment including the pass-through portion, detecting a change in the hand grip on the housing physically coupled to the display generation component; In response to detecting the change in the hand grip on the housing physically coupled to the display generation component; According to a determination that the change in the hand grip on the housing physically coupled to the display generation component meets a first criterion, replacing the first view of the three-dimensional environment with a second view of the three-dimensional environment, wherein the second view replaces at least a part of the pass-through portion with virtual content; A method comprising the above steps.
74. The first view includes a first set of virtual objects covering a first viewing angle or a range of viewing depths in front of the user's eyes; The method according to claim 73, wherein the second view includes a second set of virtual objects covering a second viewing angle or a range of viewing depths greater than the range of the first viewing angle or the viewing depth.
75. The first view includes first virtual content overlapping a first surface in the three-dimensional environment corresponding to a first physical object in the real world surrounding the user, and the second view includes, in addition to the first virtual content overlapping the first surface, second virtual content overlapping a second surface in the three-dimensional environment corresponding to a second physical object in the real world surrounding the user. The method according to claim 73.
76. The method according to claim 73, wherein the first view includes first virtual content, and the second view includes the first virtual content and second virtual content replacing the pass-through portion.
77. The method according to any one of claims 73 to 76, wherein the second view includes one or more selectable virtual objects representing one or more applications and virtual experiences respectively.
78. While displaying the second view of the three-dimensional environment, detecting a second change in the hand grip on the housing physically coupled to the display generation component; In response to detecting the second change in the hand grip on the housing physically coupled to the display generation component; In accordance with a determination that the change in the hand grip on the housing physically coupled to the display generation component meets a second criterion, replacing the second view of the three-dimensional environment with a third view of the three-dimensional environment that does not include a pass-through portion; The method according to any one of claims 73 to 77, further comprising. **Claim 79** While displaying each view of the three-dimensional environment that does not include the pass-through portion showing at least a part of the real world surrounding the user, detecting user input on the housing physically coupled to the display generation component; In response to detecting the user input on the housing physically coupled to the display generation component, redisplaying the first view including the pass-through portion including a representation of at least a part of the real world through the display generation component in accordance with a determination that the user input meets a third criterion; The method according to any one of claims 73 to 78, further comprising. **Claim 80** In response to detecting the change in the hand grip on the housing of the display generation component, Maintaining the first view of the three-dimensional environment in accordance with a determination that the change in the hand grip does not meet the first criterion; While displaying the first view of the three-dimensional environment, detecting different user input that is different from the change in the hand grip on the housing physically coupled to the display generation component, the user input causing activation of a first input device of the electronic device; In response to detecting the user input that causes activation of the first input device of the electronic device, replacing the first view of the three-dimensional environment with the second view of the three-dimensional environment; The method according to any one of claims 73 to 79, further comprising. **Claim 81** A computer-readable storage medium storing executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to execute the method according to any one of claims 73 to 80. **Claim 82** A computer system, One or more processors, A display generation component, A memory storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 73 to 80, and a memory, A computer system comprising.
83. A computer system comprising one or more processors and a display generation component, A computer system comprising means for performing the method according to any one of claims 73 to 80.
84. An information processing apparatus for use in a computer system comprising one or more processors and a display generation component, An information processing apparatus comprising means for performing the method according to any one of claims 73 to 80.
85. A method comprising: In a computer system including a display generation component and one or more input devices, Displaying a view of a virtual environment via the display generation component; Detecting a first movement of the user within the physical environment while the view of the virtual environment is being displayed and while the view of the virtual environment does not include a visual representation of a first portion of a first physical object present within the physical environment in which the user is located; In response to detecting the first movement of the user within the physical environment, The first physical object has a range that can be visible to the user based on the user's field of view with respect to the virtual environment, and according to a determination that the user is within a threshold distance of the first portion of the first physical object, changing the appearance of the view of the virtual environment in a first way that shows a physical characteristic of the first portion of the first physical object without changing the appearance of a second portion of the range of the first physical object that can be visible to the user based on the user's field of view with respect to the virtual environment; According to a determination that the user is not within the threshold distance of the first physical object present within the physical environment surrounding the user, preventing the appearance of the view of the virtual environment from being changed in the first way that shows the physical characteristic of the first portion of the first physical object; A method comprising.
86. Changing the appearance of the view of the virtual environment in the first method that indicates the physical characteristics of the first part of the first physical object, The method according to claim 85, further comprising maintaining the appearance of the view of the virtual environment in a first part of the virtual environment while changing the appearance of the view of the virtual environment in a second part of the virtual environment, wherein a boundary between the first part of the virtual environment and the second part of the virtual environment in the changed view of the virtual environment corresponds to a physical boundary of the first part of the first physical object. **Claim 87** Detecting a second movement of the user with respect to the first physical object within the physical environment, In response to detecting the second movement of the user with respect to the first physical object within the physical environment, Changing the appearance of the view of the virtual environment in a second method that indicates the physical characteristics of the second part of the first physical object according to a determination that the user is within a threshold distance of the second part of the first physical object, which is part of the range of the first physical object that may be visible to the user based on the field of view of the user with respect to the virtual environment, The method according to claim 85 or 86, further comprising. **Claim 88** The method according to any one of claims 85 to 87, wherein changing the appearance of the view of the virtual environment in the first method that indicates the physical characteristics of the first part of the first physical object further comprises displaying a translucent visual representation of the first part of the first physical object within the view of the virtual environment. **Claim 89** The method according to any one of claims 85 to 87, wherein changing the appearance of the view of the virtual environment in the first method that indicates the physical characteristics of the first part of the first physical object further comprises distorting a part of the virtual environment in a shape that represents the shape of the first part of the first physical object. **Claim 90** Changing the appearance of the view of the virtual environment in the first method that indicates the physical property of the first part of the first physical object further includes displaying a predetermined distortion of a part of the view of the virtual environment corresponding to the first part of the first physical object. The method according to any one of claims 85 to 87.
91. Detecting continuous movement of the user within the physical environment after the first movement; In response to detecting the continuous movement of the user within the physical environment after the first movement, according to the determination that the user stays within the threshold distance of the first part of the first physical object, According to the determination that the distance between the user and the first part of the first physical object has increased as a result of the continuous movement of the user within the physical environment, reducing the first display characteristic of the visual effect currently applied to the view of the virtual environment indicating the physical property of the first part of the first physical object; According to the determination that the distance between the user and the first part of the first physical object has decreased as a result of the continuous movement of the user within the physical environment, increasing the first display characteristic of the visual effect currently applied to the view of the virtual environment indicating the physical property of the first part of the first physical object; The method according to any one of claims 85 to 90, including.
92. Detecting continuous movement of the user within the physical environment after the first movement; In response to detecting the continuous movement of the user within the physical environment after the first movement, while the first physical object can be visible to the user based on the user's field of view with respect to the virtual environment, According to the determination that the distance between the user and the first part of the first physical object has increased beyond the threshold distance as a result of the continuous movement of the user within the physical environment, and according to the determination that as a result of the continuous movement of the user within the physical environment, the distance between the user and the second part of the first physical object has decreased below the threshold distance. Stopping changing the appearance of the view of the virtual environment in the first method of showing the physical characteristics of the first part of the first physical object without changing the appearance of the view of the virtual environment so as to show a second part of the first physical object that is part of the range of the first physical object that may be visible to the user based on the field of view of the user with respect to the virtual environment; Changing the appearance of the view of the virtual environment in a second method of showing the physical characteristics of the second part of the first physical object without changing the appearance of the view of the virtual environment so as to show the first part of the first physical object that is part of the range of the first physical object that may be visible to the user based on the field of view of the user with respect to the virtual environment; The method according to any one of claims 85 to 91, comprising:
93. Further comprising changing the speed of changing the appearance of the view of the virtual environment that shows the physical characteristics of the first part of the first physical object according to the speed of the first movement of the user with respect to the first physical object in the physical environment, the method according to any one of claims 85 to 92.
94. Further comprising continuously displaying a representation of at least a part of the second physical object in the view of the virtual environment that shows the physical characteristics of the second physical object, wherein the second physical object is selected by the user, the method according to any one of claims 85 to 93.
95. After changing the appearance of the view of the virtual environment in the first method of showing the physical characteristics of the first part of the first physical object, detecting a change in the posture of the user in the physical environment; In response to detecting the change in the posture and according to a determination that the change in the posture meets a first predetermined posture criterion, changing the appearance of the view of the virtual environment in each method of enhancing the visibility of the first physical object in the view of the virtual environment; The method according to any one of claims 85 to 94, further comprising:
96. Changing the appearance of the view of the virtual environment in each method of enhancing the visibility of the first physical object in the view of the virtual environment, The method according to claim 95, comprising increasing display characteristics of a visual effect currently applied to a part of the virtual environment that shows the physical characteristics of the first part of the first physical object. **Claim 97** Changing the appearance of the view of the virtual environment in each method of enhancing the visibility of the first physical object in the view of the virtual environment, The method according to claim 95, comprising changing the appearance to show the physical characteristics of an additional part of the first physical object and increasing the range of the view of the virtual environment. **Claim 98** Changing the appearance of the view of the virtual environment in each method of enhancing the visibility of the first physical object in the view of the virtual environment, The method according to claim 95, comprising changing the appearance to show the physical characteristics of all physical objects that may be visible to the user based on the user's field of view with respect to the virtual environment and increasing the range of the view of the virtual environment. **Claim 99** After detecting the change in the posture that satisfies the first predetermined posture criterion, detecting a reverse change in the posture of the user in the physical environment; In response to detecting the reverse change in the posture and according to the determination that the reverse change in the posture satisfies a second predetermined posture criterion, changing the appearance of the view of the virtual environment in each method of reversing the enhanced visibility of the first physical object in the view of the virtual environment; The method according to claim 95, further comprising. **Claim 100** After changing the appearance of the view of the virtual environment by the first method that shows the physical characteristics of the first part of the first physical object, according to the determination that a virtual field-of-view restoration criterion that requires that the positions of the user and the first part of the first physical object do not change for a first threshold time is satisfied, further reversing the change in the appearance of the view of the virtual environment in the first method. The method according to any one of claims 85 to 99. **Claim 101** A computer-readable storage medium storing executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to execute the method according to any one of claims 85 to 100.
102. A computer system comprising: one or more processors; a display generation component; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for executing the method according to any one of claims 85 to 100; A computer system comprising the above.
103. A computer system comprising one or more processors and a display generation component, The computer system comprising means for executing the method according to any one of claims 85 to 100.
104. An information processing apparatus for use in a computer system comprising one or more processors and a display generation component, The information processing apparatus comprising means for executing the method according to any one of claims 85 to 100.
Citation Information
Patent Citations
System and method for transitioning between transmissive and non-transmissive modes in a head-mounted display
JP2016532178A
Sensory feedback system and method for guiding a user in a virtual reality environment
JP2017535901A
System and method for generating a progressive representation associated with surjectively mapped virtual and physical reality image data
WO2017180990A1
Physical boundary guardian
WO2019067470A1
Image processing device, image processing method, and program
WO2019123729A1