Device, method, and graphical user interface for processing inputs to three-dimensional environment
By detecting and processing concurrent inputs from different input manipulators in a computer system, the inefficiency of existing technologies is solved, enabling more efficient and intuitive user interaction and energy savings, thus improving the user experience of augmented reality and virtual reality systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-09-19
- Publication Date
- 2026-04-17
AI Technical Summary
Existing input methods and interfaces for processing 3D environments are inefficient, complex, and error-prone in augmented reality and virtual reality systems, leading to cognitive burden on users and wasted energy in computer systems.
By detecting and processing concurrent inputs from different input manipulators in a computer system, and by utilizing visual feedback and conditional fulfillment to automatically execute operations, the number and nature of inputs are reduced, providing a more efficient human-computer interface.
It improves the seamlessness and intuitiveness of user interaction, reduces input errors, saves energy consumption of computer systems, extends battery life, and improves device operating efficiency and user experience.
Smart Images

Figure CN121889757A_ABST
Abstract
Description
[0001] Related applications This application is a continuation to and claims priority to U.S. Patent Application No. 18 / 886,891, filed September 16, 2024, which claims priority to U.S. Provisional Application No. 63 / 541,759, filed September 29, 2023. Technical Field
[0002] This disclosure relates in its entirety to computer systems that communicate with display generation components and one or more input devices and provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Technology
[0003] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices used in computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with the virtual / augmented reality environment. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention
[0004] Some methods and interfaces for processing input to 3D environments (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide incomplete support for receiving input performed using multiple input manipulators are complex, tedious, and error-prone, placing a significant cognitive burden on the user and detrimental to the experience of the virtual / augmented reality environment. Furthermore, these methods take longer than necessary, thus wasting the computer system's energy. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, there is a need for computer systems with improved methods and interfaces for processing input to computer-generated experiences, making interaction with computer systems that support multiple input manipulators more seamless, efficient, and intuitive for users. Such methods and interfaces can optionally supplement or replace conventional methods for processing input when providing extended reality experiences to users. By helping users understand the relationship between the input provided and the device's response to those inputs, such methods and interfaces reduce the quantity, extent, and / or nature of user input, thus creating a more efficient human-computer interface.
[0006] The disclosed system reduces or eliminates the aforementioned defects and other problems associated with the user interface of a computer system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, tablet, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to display generation components, the computer system also has one or more output devices, including one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, and a program or instruction set stored in memory for performing multiple functions. In some implementations, users interact with the GUI through touch and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some implementations, functions performed through interaction may optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions may optionally be included in transient and / or non-transitory computer-readable storage media or other computer program products configured for execution by one or more processors.
[0007] Electronic devices with improved methods and interfaces are needed to process input to 3D environments. Such methods and interfaces can complement or replace conventional methods for processing input to 3D environments. These methods and interfaces reduce the quantity, extent, and / or nature of user input, resulting in more efficient human-machine interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging.
[0008] According to some embodiments, a method is performed at a computer system communicating with a display generation component and one or more input devices. The method includes: displaying a user interface comprising one or more user interface objects when a view of the environment is visible via the display generation component. The method includes: detecting one or more inputs via the one or more input devices. The method includes: based on determining that the one or more inputs include a first input executed using a first input manipulator and a second input executed using a second input manipulator different from the first input manipulator, wherein the first input and the second input satisfy a concurrency criterion: providing a first input event for the first input, the first input event including information identifying a target location; and providing a second input event for the second input, the second input event including information identifying the target location.
[0009] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive; in particular, many additional features and advantages will be apparent to those skilled in the art from the accompanying drawings, description, and claims. Furthermore, it should be pointed out that the language used in this specification has been chosen in principle for readability and instruction purposes, and such choice may not be necessary to depict or define the subject matter of the invention. Attached Figure Description
[0010] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.
[0011] Figure 1A This is a block diagram illustrating the operating environment of a computer system for providing extended reality (XR) experiences according to some implementation schemes.
[0012] Figures 1B to 1P It is used in Figure 1A Examples of computer systems that provide XR experiences in the operating environment.
[0013] Figure 2 This is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience, according to some implementation schemes.
[0014] Figure 3 This is a block diagram illustrating a display generation component of a computer system configured to provide an XR experience to a user, according to some implementation schemes.
[0015] Figure 4 This is a block diagram illustrating a hand tracking unit of a computer system configured to capture user gesture input according to some implementation schemes.
[0016] Figure 5 This is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input according to some implementation schemes.
[0017] Figure 6 This is a flowchart illustrating a flare-assisted gaze tracking pipeline according to some implementation schemes.
[0018] Figures 7A to 7S Example techniques for processing inputs performed using different numbers of input manipulators, according to some implementation schemes, are illustrated.
[0019] Figures 8A to 8D It is a flowchart of a method for processing inputs using different numbers of input manipulators according to various implementation schemes. Detailed Implementation
[0020] According to some implementations, this disclosure relates to a user interface for providing extended reality (XR) experiences to users.
[0021] The systems, methods, and GUIs described in this paper improve user interface interactions with virtual / augmented reality environments in a variety of ways.
[0022] In some implementations, the computer system detects one or more inputs, and if the one or more inputs include a first input performed using a first input manipulator and a different second input performed using a different second input manipulator, and the first and second inputs meet a concurrency criterion, the computer system provides an application associated with a common target location of the inputs with a first input event for the first input and a second input event for the second input. Both the first and second input events include information identifying the target location. In some implementations, the first and second input events are associated with and optionally include information identifying different corresponding input locations other than the common target location. This allows the receiving application to distinguish concurrent inputs from different input manipulators, and also to understand the relative positions and relative movements of the input manipulators, and to use this information when performing user interface operations in response, rather than simply knowing the positions and movements of input manipulators that are independent of each other, thereby increasing the flexibility of input processing and the range of supported inputs.
[0023] Figures 1A to 6 Descriptions of example computer systems for providing XR experiences to users are provided (such as those described below with respect to method 800). Figures 7A to 7S Example techniques for processing inputs performed using different numbers of input manipulators, according to some implementation schemes, are illustrated. Figures 8A to 8DIt is a flowchart of a method for processing inputs using different numbers of input manipulators according to various implementation schemes. Figures 7A to 7S The user interface in the example is used to demonstrate Figures 8A to 8D The process in.
[0024] The processes described below enhance device operability and (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device) make the user-device interface more efficient through various technologies. These include providing users with improved visual feedback, reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional display controls, performing operations without further user input when a set of conditions has been met, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while saving storage space, and / or additional technologies. These technologies also reduce power consumption and extend device battery life by enabling users to use the device faster and more efficiently. This saves battery power and, therefore, weight, improving the device's ergonomics. These technologies also enable real-time communication, allowing the use of fewer and / or less precise sensors, resulting in more compact, lighter, and cheaper devices, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption, thereby reducing the heat generated by the device. This is especially important for wearable devices, where if a device generates too much heat, even when operating entirely within the parameters of its components, it can become uncomfortable for the user to wear.
[0025] Furthermore, in a method described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the described method can be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if the method requires performing a first step (if the conditions are satisfied) and a second step (if the conditions are not satisfied), those skilled in the art will know that the stated steps are repeated until both conditions are satisfied and conditions are not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to a method having discretionary steps, a system or computer-readable storage medium can repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0026] In some implementation schemes, such as Figure 1A As shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted display (HMD), a monitor, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted or handheld device).
[0027] In describing XR experiences, various terms are used to distinguish several related but different environments that a user can sense and / or interact with (e.g., interacting with inputs detected by the computer system 101 that generates the XR experience, causing the computer system to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms: Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0028] Extended Reality: Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the properties of virtual objects in the XR environment can be done in response to a representation of physical motion (e.g., a voice command). People can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. For example, audio objects can enable audio transparency, which selectively introduces ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, people can sense and / or interact only with audio objects.
[0029] Examples of XR include virtual reality and mixed reality.
[0030] Virtual Reality: A virtual reality (VR) environment is a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of a person's presence within the computer-generated environment and / or through the simulation of a subset of a person's physical movements within the computer-generated environment.
[0031] Mixed Reality: Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments refer to simulated environments designed to incorporate sensory input from the physical environment, or its representations, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between, but not limited to, a purely physical environment as one end and a virtual reality environment as the other. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present an MR environment can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical objects or their representations from the physical environment). For example, a system can cause movement so that virtual trees appear stationary relative to the physical ground.
[0032] Examples of mixed reality include augmented reality and augmented virtual reality.
[0033] Augmented Reality (AR): An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on a physical environment or a representation of the physical environment. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceives the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto the physical environment, such as as a hologram or onto a physical surface, allowing a person to perceive the virtual objects superimposed on the physical environment. Augmented reality environments also refer to simulated environments in which the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system can transform one or more sensor images to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. As another example, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portions can be representative but not realistic versions of the original captured image. Furthermore, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.
[0034] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment combines one or more sensory inputs from a physical environment. Sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could adopt shadows that correspond to the sun's position within the physical environment.
[0035] In augmented reality, mixed reality, or virtual reality environments, a view of the three-dimensional environment is visible to the user. This view is typically visible to the user via a virtual viewport through one or more display generating components (e.g., a display providing stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generating components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). The viewport and viewport boundary typically move with the movement of one or more display generating components (e.g., with the user's head for head-mounted devices, or with the user's hand for handheld devices such as tablets or smartphones). The user's viewpoint determines what is visible within the viewport. The viewpoint typically specifies the position and orientation relative to the 3D environment, and as the viewpoint shifts, the view of the 3D environment also shifts within the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the 3D environment that offers an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint shifts with the movement of the handheld or fixed device and / or with changes in the user's positioning relative to the handheld or fixed device (e.g., the user moves towards, away from, up, down, right, and / or left). For devices including display generation components with virtual pass-through, a portion of the physical environment visible (e.g., displayed and / or projected) via one or more display generation components is based on the field of view of one or more cameras communicating with the display generation components, which typically move with the movement of the display generation components (e.g., for head-mounted devices, moving with the movement of the user's head, or for handheld devices such as tablets or smartphones, moving with the movement of the user's hands), because the user's viewpoint moves with the movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on the movement of the user's viewpoint)).For a display generating component with optical transparency, portions of the physical environment visible through one or more display generating components (e.g., optically visible through one or more portions or fully transparent portions of the display generating component) are based on the user's field of view through the portion or fully transparent portion of the display generating component (e.g., for a head-mounted device, it moves with the movement of the user's head, or for a handheld device such as a tablet or smartphone, it moves with the movement of the user's hand), because the user's viewpoint moves with the movement of the user's field of view through the portion or fully transparent portion of the display generating component (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0036] In some embodiments, the representation of the physical environment (e.g., displayed via virtual passthrough or optical passthrough) may be partially or completely occluded by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level may optionally result in more virtual environment being displayed, replacing and / or occluding more physical environment, and decreasing the immersion level may optionally result in less virtual environment being displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, displayed with increased transparency), and one or more third background objects are de-emphasized. In some embodiments, the level of immersion includes the associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by the computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment. Optionally, this includes the number of items in the displayed background content and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generating component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view occupied by the virtual content displayed via the display generating component (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included within a background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some implementations, background content includes user interfaces (e.g., user interfaces corresponding to applications generated by a computer system), virtual objects not associated with or included in the virtual environment and / or virtual content (e.g., files generated by the computer system or representations of other users), and / or real objects (e.g., transparent objects representing real objects in the user's surrounding physical environment, visible such that they are displayed via display generation components and / or via transparent or semi-transparent components of the display generation components, because the computer system does not obscure / impede their visibility through the display generation components). In some implementations, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level may optionally be displayed concurrently with background content, which may optionally be displayed at full brightness, color, and / or semi-transparency.In some implementations, at higher immersion levels (e.g., a second immersion level above the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without concurrently displaying background content (e.g., in full-screen or fully immersive mode). Alternatively, a virtual environment displayed at a medium immersion level is displayed concurrently with darkened, blurred, or otherwise de-emphasized background content. In some implementations, the visual characteristics of background objects differ between background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are de-emphasized. In some implementations, zero immersion or a zero immersion level corresponds to a virtual environment that is de-emphasized, and instead, a representation of the physical environment (optionally having one or more virtual objects, such as an application, window, or virtual 3D object) is displayed without being occluded by the virtual environment. Using physical input elements to adjust immersion levels provides a quick and efficient way to adjust immersion, which enhances the operability of computer systems and makes user-device interfaces more efficient.
[0037] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same location and / or position within the user's viewpoint, the virtual object remains viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the direction in front of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); therefore, the user's viewpoint remains fixed even if the user's gaze shifts without moving the user's head. In embodiments where the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generating component. For example, a viewpoint-locked virtual object displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint, even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of a viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an implementation where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so the virtual object is also referred to as a "head-locked virtual object".
[0038] Environment-locked visual objects: When a computer system displays a virtual object at a location and / or position within the user's viewpoint, the virtual object is environment-locked (or, "world-locked"), the location and / or position being based on a location and / or object in a three-dimensional environment (e.g., a physical or virtual environment) (e.g., selected and / or anchored to that location and / or object with reference to it). As the user's viewpoint shifts, the location and / or object in the environment relative to the user's viewpoint changes, causing the environment-locked virtual object to appear at different locations and / or positions within the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user appears at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) so that the tree is now centered to the left in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree appears centered to the left in the user's viewpoint. In other words, the position and / or orientation of an environment-locked virtual object displayed in the user's viewpoint depends on the position to which the virtual object is locked and / or the orientation and / or orientation of the object within the environment. In some implementations, the computer system uses a stationary frame of reference (e.g., a coordinate system anchored to a fixed position and / or object in the physical environment) to determine the position of the environment-locked virtual object displayed in the user's viewpoint. An environment-locked virtual object may be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object), or it may be locked to a movable part of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot), causing the virtual object to move with the viewpoint or that part of the environment to maintain a fixed relationship between the virtual object and that part of the environment.
[0039] In some implementations, environment-locked or viewpoint-locked virtual objects exhibit lazy following behavior, reducing or delaying their movement relative to the movement of a reference point they are following. In some implementations, when exhibiting lazy following behavior, the computer system intentionally delays the movement of the virtual object when movement of the reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) is detected. For example, when the reference point (e.g., that portion of the environment or the viewpoint) moves at a first rate, the virtual object is moved by the device to remain locked to the reference point, but moves at a second rate that is slower than the first rate (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up). In some implementations, when the virtual object exhibits lazy following behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold amount, such as 0 to 5 degrees or 0 cm to 50 cm). For example, when the reference point (e.g., a portion of the environment or viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a portion of the viewpoint or environment to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a portion of the viewpoint or environment to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases to above a threshold (e.g., a "lazy following" threshold), because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some implementations, maintaining a substantially fixed position of the virtual object relative to a reference point includes displaying the virtual object within a threshold distance (e.g., 1cm, 2cm, 3cm, 5cm, 15cm, 20cm, 50cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the reference point).
[0040] Hardware: Many different types of electronic systems enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection techniques that project graphic images onto a person's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as as holograms or onto a physical surface. In some embodiments, controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 2The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within scene 105. Alternatively, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, the controller 110 is communicatively coupled to display generation components 120 (e.g., an HMD, a monitor, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., physical enclosure) of display generation component 120 (e.g., HMD or portable electronic device including a display and one or more processors), one or more input devices in input device 125, one or more output devices in output device 155, one or more sensors in sensor 190 and / or one or more peripheral devices in peripheral device 195, or shares the same physical housing or support structure with one or more of the aforementioned devices.
[0041] In some implementations, display generation component 120 is configured to provide an XR experience to a user (e.g., at least the visual component of the XR experience). In some implementations, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 3 The display generation component 120 is described in more detail. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation component 120.
[0042] According to some implementation schemes, when a user is virtually and / or physically present within scene 105, display generation component 120 provides the user with an XR experience.
[0043] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device having a display facing the user's field of view and a camera facing scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, wherein the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware used for displaying XR content (e.g., a handheld device or a tripod-mounted device) can be implemented on another type of hardware used for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interaction with XR content triggered by an interaction occurring in the space in front of a handheld device or tripod-mounted device can be similarly implemented using an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface illustrating interaction with XR content triggered by movement of a handheld device or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented using an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).
[0044] Despite Figure 1A The relevant features of the operating environment 100 are illustrated herein, but those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the exemplary embodiments disclosed herein.
[0045] Figures 1A to 1PVarious examples of computer systems for performing methods and providing audio, visual, and / or haptic feedback as part of the user interface described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first and second display assemblies 1-120a, 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b) for displaying virtual elements and / or representations of the physical environment to a user of the computer system, the virtual elements and / or the representations of the physical environment optionally being generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216 to make it easier for a user who would otherwise use glasses or contact lenses to correct their vision to view the user interface, the one or more corrective lenses optionally being removably attached to one or more optical modules in the optical modules. While many user interfaces illustrated herein show a single view of the user interface, the user interface in an HMD may optionally use two optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for display, one optical module for the user's right eye and a different optical module for the user's left eye, presenting slightly different images to the two different eyes to generate the illusion of stereoscopic depth. A single view of the user interface is typically a right-eye view or a left-eye view; the depth effect is explained in text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying status information of the computer system to the user of the computer system (when the computer system is not worn) and / or to others near the computer system. This status information may optionally be generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which may optionally be generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assemblies 1-356 and / or sensor assemblies 1-356) for detecting information about the physical environment of the device. Figure 1I One or more sensors), which can be used (optionally with one or more illuminators, such as Figure 1IThe system combines the illuminators described herein to generate digital pass-through images, capture visual media (e.g., photographs and / or videos) corresponding to the physical environment, or determine the pose (e.g., positioning and / or orientation) of physical objects and / or surfaces in the physical environment, enabling the placement of virtual objects based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assemblies 1-356 and / or...) for detecting hand positioning and / or movement. Figure 1I One or more sensors), which can be used (optionally with one or more illuminators, such as Figure 1I The illuminators 6-124 described herein (in combination) determine when one or more air gestures are performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I (Eye-tracking and gaze-tracking sensors in the system), these sensors can be used (optionally combined with one or more lights, such as...) Figure 10The light (11.3.2-110) in the image determines attention or gaze localization and / or gaze movement, which can optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for real-time communication sessions, wherein the avatar has facial expressions, hand movements, and / or body movements detected by the user based on or similar to the device. Gaze and / or attention information may optionally be combined with hand tracking information to determine user interaction with one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132 and / or dials or buttons 1-328), knobs (e.g., first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), touchpads, touchscreens, keyboards, mice and / or other input devices. One or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328) may optionally be used to perform system operations, such as recentering content in the user-visible 3D environment of the device, displaying the main user interface for launching an application, initiating a real-time communication session, or initiating the display of a virtual 3D background. A knob or digital crown (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114, and / or dials or buttons 1-328) may optionally be rotatable to adjust parameters of the visual content, such as the level of immersion of the virtual 3D environment (e.g., the extent to which the virtual content occupies the user's viewport in the 3D environment) or other parameters associated with the 3D environment and the virtual content displayed via optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).
[0046] Figure 1BExamples of head-mounted display (HMD) devices 1-100 configured to be worn by a user and provide virtual and altered / mixed reality (VR / AR) experiences are illustrated in front, top, and perspective views. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured at either end to the electronic strip assembly 1-104. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.
[0047] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of the user's head and a second strap 1-117 configured to extend above the top of the user's head. As shown, the second strap may extend between the first electronic strip 1-105a and the second electronic strip 1-105b of the electronic strip assembly 1-104. The strip assembly 1-104 and the strap assembly 1-106 may be part of a fixing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.
[0048] In at least one example, the fixing mechanism includes a first electronic strip 1-105a, which includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., the housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite to the first proximal end 1-134. The fixing mechanism may also include a second electronic strip 1-105b, which includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite to the second proximal end 1-138. The fixing mechanism may also include a first strip 1-116 and a second strip 1-117, the first strip including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strip extending between the first electronic strip 1-105a and the second electronic strip 1-105b. Strips 1-105a to b and strip 1-116 may be coupled via a connecting mechanism or assembly 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between a first proximal end 1-134 and a first distal end 1-136, and a second end 1-148 coupled to the second electronic strip 1-105b between a second proximal end 1-138 and a second distal end 1-140.
[0049] In at least one example, the first electronic strip and the second electronic strips 1-105a to b comprise plastic, metal, or other structural materials forming the shape of the substantially rigid strips 1-105a to b. In at least one example, the first strip and the second strips 1-116, 1-117 are formed of an elastic flexible material (including woven textiles, rubber, etc.). The first strip 1-116 and the second strip 1-117 may be flexible enough to conform to the shape of the user's head when wearing the HMD 1-100.
[0050] In at least one example, one or more of the first and second electronic stripes 1-105a to b may define an inner strip volume and include one or more electronic components disposed within the inner strip volume. In one example, such as Figure 1B As shown, the first electronic strip 1-105a may include electronic components 1-112. In one example, electronic components 1-112 may include a speaker. In another example, electronic components 1-112 may include computing components, such as a processor.
[0051] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is located in... Figure 1B The section marked 1-152 with dashed lines is because the display assembly 1-108 is configured to obscure the first opening 1-152 when the HMD 1-100 is assembled, as viewed from above. The housing 1-150 may also define a rearward second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover disposed in or across the front opening 1-152 to obscure the front opening 1-152 and a display screen (shown in other figures). In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 can be bent as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, wherein the display unit 1-102 is pressed.
[0052] In at least one example, the housing 1-150 may define a first hole 1-126 between a first opening 1-152 and a second opening 1-154, and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 are pressable through their respective holes 1-126 and 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a rotary dial and a pressable button. In at least one example, the first button 1-128 is a pressable and rotary dial button, and the second button 1-132 is a pressable button.
[0053] Figure 1C A rear perspective view of HMD 1-100 is illustrated. HMD 1-100 may include a light seal 1-110 extending rearwardly around the periphery of housing 1-150 of display assembly 1-108, as shown. The light seal 1-110 may be configured to extend from housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b, which are disposed at or within and / or disposed in the internal volume of housing 1-150 within a rearward-facing second opening 1-154 defined by housing 1-150 and are configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a to b may include a corresponding display screen 1-122a, 1-122b, which are configured to project light toward the user's eyes in a rearward direction through the second opening 1-154.
[0054] In at least one example, reference Figure 1B and Figure 1C Both, the display assembly 1-108 can be a front-facing display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a to b can be configured to project light in a second rearward direction opposite to the first direction. As mentioned above, the light seal 1-110 can be configured to block light from outside the HMD 1-100 from reaching the user's eyes, including by means of... Figure 1BThe front perspective view shows the light projected by the front display screen of the display assembly 1-108. In at least one example, the HMD 1-100 may also include a curtain 1-124 that blocks the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a to b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.
[0055] Figure 1B and Figure 1C Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1D to 1F In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1D to 1F Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.
[0056] Figure 1D An exploded view of an example HMD 1-200 including its various parts or components, which are separated according to the modularity and selective coupling of these components. For example, HMD 1-200 may include a strip 1-216 that may be selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first strip and the second strips 1-205a to 1-205b are removably coupled to a display unit 1-202.
[0057] Furthermore, HMD 1-200 may include a light-sealing member 1-210 configured to be removably coupled to display unit 1-202. HMD 1-200 may also include a lens 1-218, which may be removably coupled to display unit 1-202, for example, on a first display assembly including a display screen and a second display assembly. Lens 1-218 may include a custom prescription lens configured for vision correction. As noted, in Figure 1DThe exploded view shows that each component described above can be removably coupled, attached, reattached, and replaced to update the component, or replaced for different users. For example, belts such as belt 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a to b can be replaced according to the user, so that these parts are customized to fit and correspond to the individual user of HMD 1-200.
[0058] Figure 1D Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1B , Figure 1C and Figures 1E to 1F In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1B , Figure 1C and Figures 1E to 1F Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1D Examples of devices, features, components, and parts are shown.
[0059] Figure 1E An exploded view illustrating an example of a display unit 1-306 of an HMD is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0060] In at least one example, the display unit 1-306 may further include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a to b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a to b has at least one motor, such that the motor is capable of translating the display screens 1-322a to b to match the interpupillary distance of the user's eyes.
[0061] In at least one example, display unit 1-306 may include a dial or button 1-328 that is pressable relative to frame 1-350 and accessible to a user outside frame 1-350. Button 1-328 may be electrically connected to motor assembly 1-362 via a controller, such that button 1-328 can be operated by a user to cause the motor of motor assembly 1-362 to adjust the positioning of display screens 1-322a to b.
[0062] Figure 1E Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1D and Figure 1F In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1D and Figure 1F Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1E Examples of devices, features, components, and parts are shown.
[0063] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first and second corresponding display screen for interpupillary adjustment, as described above.
[0064] References in this article Figures 1B to 1E The following figures, which are referenced in this disclosure, will be used to describe the subject in more detail. Figure 1F The exploded view shows the various parts, systems, and assemblies. Figure 1F The display unit 1-406 shown can be connected with Figures 1B to 1E The shown fixture assembly and integration includes electronic strips, belts, and other components (including light seals, connecting assemblies, etc.).
[0065] Figure 1F Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1EIn any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1E Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1F Examples of devices, features, components, and parts are shown.
[0066] Figure 1G An example is the front cover assembly 3-100 of the HMD device described herein (e.g., Figure 1G An exploded perspective view of the front cover assembly 3-1) of the HMD 3-100 shown or any other HMD device shown and described herein. Figure 1G The front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or “cover”), an adhesive layer 3-106, a display assembly 3-108 including a biconvex lens panel or array 3-110, and a structural decorative element 3-112. The adhesive layer 3-106 secures the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the decorative element 3-112. The decorative element 3-112 secures the various components of the front cover assembly 3-100 to the frame or base of the HMD device.
[0067] In at least one example, such as Figure 1G As shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including a biconvex lens array 3-110 can be bent to adapt to the curvature of a user's face. The transparent cover 3-102 and the shield 3-104 can be bent in two or three dimensions, for example, vertically in and out of the Z-plane along the Z direction, and horizontally in and out of the Z-plane along the X direction. In at least one example, the display assembly 3-108 may include the biconvex lens array 3-110 and a display panel with pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., the horizontal direction) to adapt to the curvature of a user's face from one side (e.g., the left) to the other (e.g., the right). In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in the following figures, but may include the biconvex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to accommodate the curvature of the user's face.
[0068] In at least one example, the cover 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the cover 3-104 may include one or more opaque portions, such as opaque ink-printed portions or other opaque film portions on the back of the cover 3-104. When the HMD device is worn, the rear surface may be the surface of the cover 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the cover 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the cover 3-104 may include peripheral portions that visually conceal any components surrounding the outer periphery of the display screen of the display assembly 3-108. In this way, the opaque portions of the cover conceal any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the cover 3-104, including electronic components, structural components, etc.
[0069] In at least one example, the housing 3-104 may define one or more transparent aperture portions 3-120 through which the sensor can transmit and receive signals. In one example, portion 3-120 is an aperture through which the sensor can extend or transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion more transparent than the surrounding translucent or opaque portion of the housing, through which the sensor can transmit and receive signals through the housing and via the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0070] Figure 1G Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, individually or in any combination. Figure 1G Examples of devices, features, components, and parts are shown.
[0071] Figure 1H An exploded view of an example HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102 comprising one or more sensors, cameras, projectors, etc., mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 on which one or more sensors of the sensor system 6-102 may be fixed / secured.
[0072] Figure 1I A portion of an HMD device 6-100, including a front transparent cover 6-104 and a sensor system 6-102, is illustrated. The sensor system 6-102 may include multiple different sensors, transmitters, and receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated on the front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter in system 6-102. As referenced herein, "side," "side," "lateral," "horizontal," and other similar terms refer to... Figure 1J The orientation or direction indicated by the X-axis. Terms such as "vertical," "upward," "downward," and similar terms refer to the orientation or direction indicated by... Figure 1J The orientation or direction indicated by the Z-axis. Terms such as "frontward," "rearward," "forward," "backward," and similar terms refer to the orientation or direction indicated by the Z-axis. Figure 1J The orientation or direction indicated by the Y-axis shown.
[0073] In at least one example, a transparent cover 6-104 may define the front outer surface of an HMD device 6-100, and a sensor system 6-102, including various sensors and their components, may be positioned behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through it, including both light detected by the sensor system 6-102 and light emitted therefrom.
[0074] As mentioned elsewhere herein, the HMD device 6-100 may include one or more controllers, which include processors for electrically coupling various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as displays. Furthermore, as will be shown in more detail below with reference to other accompanying figures, various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to the HMD device 6-100. Figure 1I Various structural frame components, brackets, etc., not shown. For clarity, Figure 1I The components of sensor system 6-102 are shown, which are not attached to or electrically coupled to other components.
[0075] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. These instructions may include, or cause the processor to execute, one or more algorithms for self-correcting the angle and position of the various cameras described herein as the camera's initial positioning, angle, or orientation is affected by collisions or deformations due to accidental drop events or other events over time.
[0076] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. System 6-102 may include two scene cameras 6-102, respectively positioned on either side of the nose bridge or arched structure of the HMD device 6-100, such that each of the two cameras 6-106 approximately corresponds to the positioning of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and, when the HMD device 6-100 is used, provide images and content for MR video pass-through to a display screen facing the user's eyes. The scene cameras 6-106 may also be used for environment and object reconstruction.
[0077] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 that is generally pointing forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction as well as user hand and body tracking. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centrally located along the width of the HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be located above the central bridge of the nose or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction as well as hand and body tracking. In at least one example, the second depth sensor may include a LiDAR sensor.
[0078] In at least one example, the sensor system 6-102 may include a depth projector 6-112, which is typically forward-facing to project electromagnetic waves (e.g., in the form of a predetermined spot pattern) into or within the field of view of the user and / or scene camera 6-106, or into or beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a spot pattern, which are reflected from an object and back into the aforementioned depth sensors, including depth sensors 6-108 and 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction, as well as hand and body tracking.
[0079] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114, whose field of view is generally directed downwards relative to the HMD device 6-100 on the Z-axis. In at least one example, the downward-facing camera 6-114 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, head-mounted device tracking, and facial avatar detection and creation for displaying a user avatar on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward-facing camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.
[0080] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, head-mounted device tracking, and facial avatar detection and creation for displaying a user avatar on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. This is used for hand and body tracking, head-mounted device tracking, and facial avatar creation. In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views along the X-axis or in a direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, head-mounted device tracking, and facial avatar detection and reconstruction.
[0081] In at least one example, the sensor system 6-102 may include multiple eye-tracking and gaze-tracking sensors for determining identity, status, and the user's gaze direction during and / or prior to use. In at least one example, the eye / gaze-tracking sensor may include a nose-eye camera 6-120 positioned on either side of the user's nose and adjacent to the user's nose when wearing the HMD device 6-100. The eye / gaze sensor may also include a bottom eye camera 6-122 positioned below the respective user's eye for capturing images of the eye for use in facial avatar detection and creation, gaze tracking, and iris identification functions.
[0082] In at least one example, sensor system 6-102 may include an infrared illuminator 6-124 that is pointed outward from HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of sensor system 6-102. In at least one example, sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, flicker sensor 6-126 may detect the refresh rate of the overhead light to avoid display flicker. In one example, infrared illuminator 6-124 may include a light-emitting diode and may be specifically designed for low-light environments to illuminate a user's hands and other objects in low light for detection by the infrared sensors of sensor system 6-102.
[0083] In at least one example, multiple sensors (including scene camera 6-106, downward camera 6-114, chin camera 6-116, side camera 6-118, depth projector 6-112, and depth sensors 6-108, 6-110) can be combined with an electrically coupled controller to combine depth data with camera data for hand tracking and for sizing, thereby improving the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the above-described and Figure 1I The downward-facing camera 6-114, the chin camera 6-116, and the side camera 6-118 shown can be wide-angle cameras capable of operating in both the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, and 6-118 can operate solely in black-and-white light detection to simplify image processing and achieve sensitivity.
[0084] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1J to 1L In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1J to 1LAny of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1I Examples of devices, features, components, and parts are shown.
[0085] Figure 1J A lower perspective view of an example HMD 6-200 including a cover or shield 6-204 fixed to a frame 6-230 is shown. In at least one example, a sensor 6-203 of a sensor system 6-202 may be disposed around the periphery of the HMD 6-200 such that the sensor 6-203 is disposed outwardly around the periphery of the display area or region 6-232 so as not to obstruct the view of the displayed light. In at least one example, the sensor may be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing light to pass back and forth through the shield 6-204 by the sensor and the projector. In at least one example, an opaque ink or other opaque material or film / layer may be disposed on the shield 6-204 around the display area 6-232 to conceal the components of the HMD 6-200 outside the display area 6-232 rather than through a transparent portion defined by the opaque portion through which the sensor and the projector transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area surrounding the periphery of the display and the shield 6-204.
[0086] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 transmits and receives signals. In the illustrated examples, the sensor 6-203 of the sensor system 6-202, which transmits and receives signals through the shield 6-204, or more specifically through the transparent area 6-209 defined by the opaque portion 6-207 of the shield 6-204, may include... Figure 1I The examples illustrate the same or similar sensors, such as depth sensors 6-108 and 6-110, depth projector 6-112, first scene camera and second scene camera 6-106, first downward camera and second downward camera 6-114, first side camera and second side camera 6-118, and first infrared illuminator and second infrared illuminator 6-124. These sensors also... Figure 1K and Figure 1LThe example is shown. Other sensors, sensor types, number of sensors, and their relative positioning can be included in one or more other examples of the HMD.
[0087] Figure 1J Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1I and Figures 1K to 1L In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1I and Figures 1K to 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1J Examples of devices, features, components, and parts are shown.
[0088] Figure 1K A front view of a portion of an example of an HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330, is shown. Figure 1K The examples shown do not include a front cover or shield to illustrate brackets 6-336 and 6-338. For example, Figure 1J The shield 6-204 shown includes an opaque portion 6-207 that visually covers / blocks the view of anything outside the display / display area 6-334 (e.g., radially / peripherally outside the display / display area), including the sensor 6-303 and the bracket 6-338.
[0089] In at least one example, various sensors of sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, scene camera 6-306 includes strict tolerances for angles relative to each other. For example, the tolerance for the mounting angle between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such strict tolerances, in one example, scene camera 6-306 may be mounted to bracket 6-338 instead of a housing. The bracket may include a cantilever on which scene camera 6-306 and other sensors of sensor system 6-302 may be mounted to maintain their positioning and orientation in the event of a drop event caused by a user that results in any deformation of other brackets 6-226, housing 6-330, and / or housing.
[0090] Figure 1K Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1J and Figure 1LIn any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1I to 1J and Figure 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1K Examples of devices, features, components, and parts are shown.
[0091] Figure 1L A bottom view illustrating an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is shown. The sensor system 6-402 is compatible with the above and other parts of this document (including references). Figures 1I to 1K Other sensor systems described are similar. In at least one example, the jaw camera 6-416 may be oriented downwards to capture images of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to a frame or housing 6-430 or one or more internal brackets that are directly coupled to the frame or housing 6-430 shown. The frame or housing 6-430 may include one or more holes / openings 6-415 through which the jaw camera 6-416 transmits and receives signals.
[0092] Figure 1L Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1K In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1I to 1K Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1L Examples of devices, features, components, and parts are shown.
[0093] Figure 1MA rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated. This IPD adjustment system includes a first optical module and a second optical module 11.1.1-104a-104a-105a-106 ... In at least one example, buttons 11.1.1-114 can be electrically communicated with the first motor and the second motors 11.1.1-110a to b via a processor or other circuit components to activate the first motor and the second motors 11.1.1-110a to b and respectively cause the first optical module and the second optical modules 11.1.1-104a to b to change their positions relative to each other.
[0094] In at least one example, the first and second optical modules 11.1.1-104a to b may include corresponding display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user can manipulate (e.g., press and / or rotate) buttons 11.1.1-114 to activate positional adjustment of the optical modules 11.1.1-104a to b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a to b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD, such that the optical modules 11.1.1-104a to b can be adjusted to match the IPD.
[0095] In one example, a user can manipulate buttons 11.1.1-114 to cause automatic positional adjustment of the first and second optical modules 11.1.1-104a to b. In another example, a user can manipulate buttons 11.1.1-114 to cause manual adjustment, moving the optical modules 11.1.1-104a to b further or closer (e.g., when the user rotates buttons 11.1.1-114 in one way or another) until the user visually matches their own IPD. In one example, manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a to b via motors 11.1.1-110a to b is supplied by a power source. In another example, the adjustment and movement of the optical modules 11.1.1-104a to b via the manipulation buttons 11.1.1-114 are mechanically actuated via the movement buttons 11.1.1-114.
[0096] Figure 1MAny of the features, components, and / or parts shown (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and parts shown and described herein in any other illustrated figures. Similarly, any of the features, components, and / or parts shown and described herein with reference to any other illustrated figures (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and / or parts shown and described herein. Figure 1M Examples of devices, features, components, and parts are shown.
[0097] Figure 1N A front perspective view of a portion of HMD 11.1.2-100 is shown, including an outer structural frame 11.1.2-102 defining first and second holes 11.1.2-106a, 11.1.2-106b, and an inner or intermediate structural frame 11.1.2-104. Holes 11.1.2-106a to b are located in... Figure 1N The holes 11.1.2-106a to b are shown in dashed lines because viewing of the HMD 11.1.2-100 may be obstructed by one or more other components coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as illustrated. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second holes 11.1.2-106a to b.
[0098] Mounting brackets 11.1.2-108 may include intermediate or central portions 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portions 11.1.2-109 may not be the geometric center or middle of the brackets 11.1.2-108. Instead, the intermediate / central portions 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portions 11.1.2-109. In at least one example, mounting bracket 108 includes first cantilever arms 11.1.2-112 and second cantilever arms 11.1.2-114 extending away from the intermediate portions 11.1.2-109 of the mounting brackets 11.1.2-108 coupled to the inner frame 11.1.2-104.
[0099] like Figure 1NAs shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to adapt to the user's nose when the user wears the HMD 11.1.2-100. This curved geometry may be referred to as the bridge of the nose 11.1.2-111 and is centrally located on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between holes 11.1.2-106a and b, such that the cantilever 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the central portion 11.1.2-109 to complement the nose bridge geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to adapt to the user's nose, as mentioned above. The geometry of the bridge of the nose 11.1.2-111 adapts to the nose, as it provides a curvature that conforms to the shape of the user's nose, offering a comfortable fit from above, above, and around.
[0100] The first cantilever 11.1.2-112 may extend in a first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108, and the second cantilever 11.1.2-114 may extend in a second direction opposite to the first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as “cantilever” or “cantilever” arms because each arm 11.1.2-112, 11.1.2-114 includes free distal ends 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, arms 11.1.2-112 and 11.1.2-114 extend from the middle section 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102 and 11.1.2-104 are not attached.
[0101] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include multiple sensors 11.1.2-110a-f. Each of the multiple sensors 11.1.2-110a-f may include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f may be used for object recognition in three-dimensional space, making it important to maintain the precise relative positioning of two or more of the multiple sensors 11.1.2-110a-f. The cantilever nature of the mounting bracket 11.1.2-108 protects the sensors 11.1.2-110a-f from damage and displacement in the event of an accidental drop by the user. Because the sensors 11.1.2-110a-f cantilevered on the arms 11.1.2-112 and 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the internal frame and / or the external frames 11.1.2-104 and 11.1.2-102 are not transmitted to the cantilever arms 11.1.2-112 and 11.1.2-114, and therefore do not affect the relative position of the sensors 11.1.2-110a-f coupled to / mounted to the mounting bracket 11.1.2-108.
[0102] Figure 1N Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in any other example of the devices, features, components, and other examples described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 1N Examples of devices, features, components, and parts are shown.
[0103] Figure 10 An example of optical modules 11.3.2-100 for use in electronic devices, such as HMDs, including the HMD devices described herein, is illustrated. As shown in one or more other examples described herein, optical modules 11.3.2-100 may be one of two optical modules within an HMD, wherein each optical module is aligned to project light toward a user's eye. In this way, a first optical module may project light toward a user's first eye via a display screen, and a second optical module of the same device may project light toward a user's second eye via another display screen.
[0104] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a tube or optical module tube. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the user's eyes when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0105] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eye during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the camera 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward the user's eyes when the HMD is worn. The individual lights 11.3.2-110 in the light strips 11.3.2-108 may be spaced apart around the light strips 11.3.2-108, and thus spaced evenly or unevenly around the display 11.3.2-104 at various locations on the light strips 11.3.2-108 and around the display 11.3.2-104.
[0106] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light onto the user's eyes through the viewing opening 11.3.2-101. In one example, a camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.
[0107] As mentioned above, Figure 10Each of the components and features of the optical modules 11.3.2-100 shown can be replicated in another (e.g., a second) optical module set up with the HMD to interact with the user’s other eye (e.g., project light and capture images).
[0108] Figure 10 Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1P Any other example of the device, feature, component, and part shown or otherwise described herein. Similarly, refer to... Figure 1P Any of the features, components, and / or parts shown, described, or otherwise described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 10 Examples of devices, features, components, and parts are shown.
[0109] Figure 1P A cross-sectional view of an example optical module 11.3.2-200 is shown, which includes a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. Channels 11.3.2-212 and 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of an HMD device to allow the optical module 11.3.2-200 to be adjusted and positioned relative to the user's eye to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guide rod to secure the optical module 11.3.2-200 in the appropriate position within the HMD.
[0110] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eye when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eye. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, lenses 11.3.2-216 are positioned above light strips 11.3.2-208 and one or more eye-tracking cameras 11.3.2-206, such that cameras 11.3.2-206 are configured to capture an image of a user's eye through lenses 11.3.2-216, and light strips 11.3.2-208 include lamps configured to project light onto the user's eye through lenses 11.3.2-216 during use.
[0111] Figure 1P Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in any of the other examples of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 1P Examples of devices, features, components, and parts are shown.
[0112] Figure 2This is a block diagram of an example controller 110 according to some implementation schemes. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the implementation schemes disclosed herein. Therefore, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0113] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0114] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 220 may optionally include one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures, or subsets thereof, including optional operating system 230 and XR experience module 240.
[0115] Operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., single XR experiences for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various embodiments, XR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.
[0116] In some implementations, the data acquisition unit 242 is configured to acquire data from... Figure 1A The data acquisition unit 242 includes at least the display generation component 120, and optionally acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 242 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0117] In some implementations, the tracking unit 244 is configured to map scene 105, and the tracking at least shows the generating component 120 relative to... Figure 1A The tracking unit 244 tracks the location / position of scene 105, and optionally tracks the position of one or more of input devices 125, output devices 155, sensors 190, and / or peripheral devices 195. To this end, in various embodiments, the tracking unit 244 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics. In some embodiments, the tracking unit 244 includes a hand tracking unit 245 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 245 is configured to track the location / position of one or more portions of the user's hand, and / or the position of one or more portions of the user's hand relative to the user's hand. Figure 1A The motion of scene 105 relative to the display generation component 120 and / or relative to a coordinate system (defined relative to the user's hand). The following refers to the motion relative to... Figure 4 The hand tracking unit 245 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the user's gaze (or more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to XR content displayed via display generation component 120. The following description is relative to... Figure 5 The eye-tracking unit 243 is described in more detail.
[0118] In some implementations, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripheral device 195. To this end, in various implementations, coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0119] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0120] Although the data acquisition unit 242, the tracking unit 244 (e.g., including eye tracking unit 243 and hand tracking unit 245), the coordination unit 246, and the data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including eye tracking unit 243 and hand tracking unit 245), the coordination unit 246, and the data transmission unit 248 may reside in a separate computing device.
[0121] also, Figure 2 This is used more as a functional description of various features that can exist in a particular specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0122] Figure 3This is a block diagram illustrating an example of generating component 120 according to some embodiments. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, and / or similar interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.
[0123] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0124] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, display generation component 120 (e.g., HMD) includes a single XR display. Alternatively, display generation component 120 may include XR displays for each of the user's eyes. In some embodiments, the one or more XR displays 312 are capable of presenting MR and VR content.
[0125] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to the scene the user would see in the absence of a display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0126] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 may optionally include one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and XR rendering module 340.
[0127] Operating system 330 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. For the aforementioned purposes, in various embodiments, XR rendering module 340 includes data acquisition unit 342, XR rendering unit 344, XR mapping generation unit 346, and data transmission unit 348.
[0128] In some implementations, the data acquisition unit 342 is configured to acquire data from at least... Figure 1A The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0129] In some implementations, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. To this end, in various implementations, the XR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0130] In some implementations, the XR mapping generation unit 346 is configured to generate XR maps based on media content data (e.g., 3D maps of mixed reality scenes or maps in which computer-generated objects can be placed to generate extended reality physical environments). To this end, in various implementations, the XR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0131] In some implementations, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various implementations, the data transmission unit 348 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0132] Although the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 are shown residing in a single device (e.g., Figure 1A The data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 are located on the display generation component 120, but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in a separate computing device.
[0133] also, Figure 3 This serves more as a functional description of various features that may exist in a particular specific implementation, and differs from the structural schematic diagram of the implementation described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 3 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0134] Figure 4 This is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) by hand tracking unit 245 ( Figure 2 Controls are used to track the location / position of one or more parts of the user's hand, and / or the location of one or more parts of the user's hand relative to the user's hand. Figure 1AThe movement is relative to scenario 105 (e.g., relative to a portion of the user's surrounding physical environment, relative to display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head)) and / or relative to a coordinate system defined relative to the user's hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0135] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures at least three-dimensional scene information including the human user's hand 406. The image sensor 404 captures hand images at a sufficient resolution to distinguish fingers and their corresponding positions. The image sensor 404 typically captures images of other parts of the user's body, or possibly all parts of the body, and may have scaling capabilities or a dedicated sensor with increased magnification to capture images of the hand at the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of scene 105, or serves as the image sensor for capturing the physical environment of scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a manner that uses the field of view of the image sensor 404 or a portion thereof to define an interaction space in which hand movements captured by the image sensor are considered input to the controller 110.
[0136] In some implementations, image sensor 404 outputs a sequence of frames containing 3D image data (and, in addition, possibly color image data) to controller 110, which extracts high-level information from the image data. This high-level information is typically provided via an application programming interface (API) to an application running on the controller, which in turn drives display generation component 120. For example, a user can interact with software running on controller 110 by moving their hand 406 and / or changing their hand gestures.
[0137] In some embodiments, image sensor 404 projects a speckle pattern onto a scene including hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) via triangulation based on the lateral displacement of the specks in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from image sensor 404. In this disclosure, it is assumed that image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand-tracking device) may use other 3D mapping methods, such as stereo imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.
[0138] In some implementations, hand tracking device 140 captures and processes time-series depth maps containing the user's hand as the user moves their hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D map data to extract image block descriptors of the hand from these depth maps. The software may match these descriptors with image block descriptors stored in database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D position of the user's hand joints and fingertips.
[0139] The software can also analyze the trajectories of the hand and / or fingers across multiple frames in a sequence to identify gestures. The pose estimation function described herein can be alternated with motion tracking, such that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find pose changes occurring in the remaining frames. Pose, motion, and gesture information is provided to an application running on controller 110 via the aforementioned API. The application can, for example, move and modify the image presented on display generation component 120 in response to the pose and / or gesture information, or perform other functions.
[0140] In some implementations, gestures include air gestures. An air gesture is a gesture detected without the user touching an input element that is part of the device (e.g., computer system 101, one or more input devices 125 and / or hand tracking device 140) (or independent of an input element that is part of the device) and based on the detected movement of a part of the user's body (e.g., head, one or two arms, one or two hands, one or more fingers and / or one or two legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., including a tapping gesture in which the hand moves a predetermined amount and / or rate in a predetermined pose, or a shaking gesture that includes a predetermined rate or amount of rotation of a part of the user's body)).
[0141] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures performed by the movement of a user's fingers relative to other fingers or portions of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, air gestures are detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and are based on the detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or portion of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or rate in a predetermined pose, or a shaking gesture that includes a predetermined rate or amount of rotation of a part of the user's body)).
[0142] In some implementations where the input gesture is an air gesture (e.g., where, in the absence of physical contact with the input device, the input device provides the computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen, or contact with a mouse or touchpad to move the cursor to the user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Therefore, in specific implementations involving air gestures, for example, the input gesture combined with (e.g., concurrently) the movement of the user's fingers and / or hand detects attention (e.g., gaze) toward a user interface element to perform pinch and / or tap input, as described in more detail below.
[0143] In some implementations, input gestures directed to a user interface object are performed, either directly or indirectly, by referencing the user interface object. For example, user input is performed directly on the user interface object based on the user's hand performing an input gesture at a location corresponding to the user interface object's position in the three-dimensional environment (e.g., determined based on the user's current viewpoint). In some implementations, when user attention to the user interface object (e.g., gazing) is detected, input gestures are performed indirectly on the user interface object based on the user's hand not being positioned at a location corresponding to the user interface object's position in the three-dimensional environment while the user is performing the input gesture. For example, for direct input gestures, the user can guide their input to the user interface object by initiating a gesture at or near a location corresponding to the user interface object's display position (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge or center of the option). For indirect input gestures, the user can guide their input to the user interface object by focusing on it (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the user interface object's display position).
[0144] In some implementations, the input gestures (e.g., air gestures) used in the various examples and implementations described herein include pinch input and tap input for interacting with a virtual or mixed reality environment. For example, the pinch input and tap input described below are executed as air gestures.
[0145] In some implementations, pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 to 1 second) interruption of contact. A long pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact is detected. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some implementations, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively with each other immediately (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts the contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).
[0146] In some embodiments, pinch-and-drag gestures as air gestures (e.g., air drag gestures or air swipe gestures) include pinch gestures (e.g., pinch gestures or long pinch gestures) performed in conjunction with (e.g., following) a drag input that changes the user's hand position from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user holds the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opening two or more of their fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together to touch each other and uses the drag gesture to move the same hand to the second position in the air). In some implementations, pinch input is performed by the user's first hand, and drag input is performed by the user's second hand (e.g., while the user continues pinch input with the user's first hand, the user's second hand moves in the air from a first position to a second position). In some implementations, input gestures as air gestures include inputs performed using both of the user's hands (e.g., pinch and / or tap inputs). For example, input gestures include two (e.g., more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., pinch input, long pinch input, or pinch and drag input) is performed using the user's first hand, and a second pinch input is performed using the other hand (e.g., the second hand in both of the user's hands). In some implementations, movement between the user's two hands is performed (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands).
[0147] In some implementations, a tap input performed as an air gesture (e.g., pointing at a user interface element) includes movement of a user's finger toward the user interface element, movement of the user's hand toward the user interface element (optionally, the user's finger extends toward the user interface element), downward movement of the user's finger (e.g., mimicking a mouse click or a tap on a touchscreen), or other predefined movements of the user's hand. In some implementations, the tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the finger or hand moving away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by the end of the movement. In some implementations, the end of the movement is detected based on changes in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of finger or hand movement, and / or a reversal of the acceleration direction of finger or hand movement).
[0148] In some implementations, the user's attention is determined to be directed to a portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment (optionally, no other conditions are required). In some implementations, the user's attention is determined to be directed to that portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment using one or more additional conditions, such as requiring the gaze to be directed to that portion of the 3D environment for at least a threshold duration (e.g., dwell time) and / or requiring the gaze to be directed to that portion of the 3D environment when the user's viewpoint is within a distance threshold from that portion of the 3D environment, so that the device determines that the user's attention is directed to that portion of the 3D environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the 3D environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0149] In some implementations, the detection of the readiness configuration of a user or a portion of a user is performed by a computer system. The detection of the hand's readiness configuration is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the readiness of the hand is determined based on whether it has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grasping gesture, or a pre-tap with one or more fingers extended and the back of the hand facing the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user above the user's waist and below the user's head, or moving away from the user's body or legs). In some implementations, the readiness state is used to determine whether an interactive element of the user interface responds to attentional (e.g., gaze) input.
[0150] In scenarios where input is described by reference to air gestures, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such gestures. Optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the spatial positioning of the hardware input device, and the positioning and / or movement of the hardware input device can be used in place of the positioning and / or movement of the one or two hands corresponding to the air gesture. Similarly, in scenarios where input is described by reference to air pose, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such poses. User input can be detected using controls contained in hardware input devices, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, a hand or finger cover that can detect the position or positional change of a portion of a hand and / or finger relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, wherein user input using controls contained in the hardware input device replaces hand and / or finger gestures such as air taps or air pinches in corresponding air gestures. For example, a selection input described as being performed using an air tap or air pinch input can alternatively be detected using button presses, taps on touch-sensitive surfaces, presses on pressure-sensitive surfaces, or other hardware inputs. As another example, motion input described as being performed using air pinch and drag (e.g., air drag gestures or air swipe gestures) can alternatively be detected based on interaction with hardware input controls (such as button press and hold, touch on a touch-sensitive surface, press on a pressure-sensitive surface, or other hardware input following movement of a hardware input device (e.g., together with a hand associated with the hardware input device) through space). Similarly, two-handed input involving movement of hands relative to each other can be performed using an air gesture and a hardware input device not in the hand performing the air gesture, two hardware input devices held in different hands, or two air gestures performed by different hands using an air gesture and / or inputs detected by one or more hardware input devices described above.
[0151] In some embodiments, the software may be downloaded to controller 110 electronically, for example, via a network, or alternatively, may be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, database 408 is also stored in memory associated with controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware, such as custom or semi-custom integrated circuits or programmable digital signal processors (DSPs). Although in Figure 4 The controller 110 is shown, but for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing function of the image sensor 404 may also be integrated into a computer or other computerized device controlled by the sensor output.
[0152] Figure 4 Also included is a schematic diagram of a depth map 410 captured by image sensor 404 according to some embodiments. As explained above, the depth map comprises a matrix of pixels with corresponding depth values. Pixel 412 corresponding to hand 406 has been segmented from the background and wrist in the map. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where gray shadows become darker as depth increases. Controller 110 processes these depth values to identify and segment components of the image that have human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and frame-to-frame motion from the depth map sequence.
[0153] Figure 4 The controller 110 also schematically illustrates, according to some embodiments, the hand skeleton 414 ultimately extracted from the depth map 410 of the hand 406. Figure 4 In this configuration, the hand skeleton 414 is superimposed on the hand background 416, which has already been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, fingertips, the center of the palm, the end of the hand connecting to the wrist, etc.), are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.
[0154] Figure 5 An eye-tracking device 130 is illustrated. Figure 1A Example implementation of ). In some implementations, the eye-tracking device 130 consists of an eye-tracking unit 243 ( Figure 2The eye-tracking device 130 is controlled to track the positioning and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as a head-mounted device, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye-tracking device 130 may optionally be a separate device from the handheld device or XR room. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 may optionally be used in conjunction with a display generation component that is also head-mounted or not head-mounted. In some embodiments, the eye-tracking device 130 is not a head-mounted device and may optionally be used in conjunction with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device and may optionally be part of a non-head-mounted display generation component.
[0155] In some embodiments, the display generation component 120 uses display mechanisms (e.g., a left near-eye display panel and a right near-eye display panel) to display frames including left and right images in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or semi-transparent display on which virtual objects are displayed, allowing the user to view the physical environment directly through the transparent or semi-transparent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects may, for example, be projected onto a physical surface or as holograms, allowing an individual to observe virtual objects superimposed on the physical environment using the system. In this case, separate display panels and image frames for the left and right eyes may not be necessary.
[0156] like Figure 5As shown, in some embodiments, eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an array or ring of IR or NIR light sources, such as LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be pointed at the user's eye to receive IR or NIR light reflected directly from the eye, or alternatively, it may be pointed at "hot" mirrors located between the user's eye and the display panel, which reflect the IR or NIR light from the eye back to the eye-tracking camera while allowing visible light to pass through. Eye-tracking device 130 may optionally capture images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps), analyze these images to generate gaze tracking information, and transmit the gaze tracking information to controller 110. In some embodiments, the user's two eyes are tracked separately using corresponding eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked using corresponding eye-tracking cameras and illumination sources.
[0157] In some implementations, a device-specific calibration procedure is used to calibrate the eye-tracking device 130 to determine parameters for the eye-tracking device in a specific operating environment 100, such as the 3D geometry and parameters of the LEDs, camera, thermal mirror (if present), eye lenses, and display. The device-specific calibration procedure can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration procedure can be automated or manual. According to some implementations, a user-specific calibration procedure may include estimations of eye parameters for a specific user, such as pupil position, foveal position, optical axis, visual axis, interocular distance, etc. According to some implementations, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, a flash-assisted method can be used to process the images captured by the eye-tracking camera to determine the current visual axis and the user's gaze point relative to the display.
[0158] like Figure 5As shown, the eye-tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system. The gaze tracking system includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye-tracking camera 540 may be pointed toward a mirror 550 located between the user's eye 592 and a display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, projector, etc.). These mirrors reflect the IR or NIR light from the eye 592 while allowing visible light to pass through. Figure 5 (as shown in the top portion), or alternatively, it can be pointed towards the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion), Figure 5 (As shown in the bottom part).
[0159] In some implementations, controller 110 renders AR or VR frames 562 (e.g., left and right frames for the left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye-tracking camera 540 for various purposes, such as processing frame 562 for display. Controller 110 may optionally estimate the user's gaze point on display 510 based on the gaze tracking input 542 obtained from eye-tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated based on gaze tracking input 542 may optionally be used to determine the direction the user is currently looking.
[0160] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, controller 110 can render virtual content differently based on the determined user gaze direction. For example, controller 110 can generate virtual content at a higher resolution in the concave region determined according to the user's current gaze direction than in the peripheral region. Alternatively, the controller can position or move virtual content in the view based at least partially on the user's current gaze direction. Also, the controller can display specific virtual content in the view based at least partially on the user's current gaze direction. As another example use case in AR applications, controller 110 can guide an external camera used to capture the physical environment of an XR experience to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently looking at on display 510. As another example use case, eye lens 520 can be a focusable lens, and the controller uses gaze tracking information to adjust the focus of eye lens 520 so that the virtual object the user is currently looking at has appropriate convergence / divergence to match the convergence of the user's eyes 592. The controller 110 can use gaze tracking information to guide the eye lens 520 to adjust its focus so that the nearby object that the user is looking at appears at the correct distance.
[0161] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lens 520), an eye-tracking camera (e.g., eye-tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each lens in the head-mounted device, such as... Figure 5 As shown. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0162] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the positions and angles of the eye-tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at different wavelengths (e.g., 940 nm) may be used on each side of the user's face.
[0163] like Figure 5 The gaze tracking system implementations illustrated herein can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0164] Figure 6 Examples of flash-assisted gaze tracking pipelines according to some embodiments are illustrated. In some embodiments, the gaze tracking pipeline uses a flash-assisted gaze tracking system (e.g., such as...) Figure 1A and Figure 5 The illustrated eye-tracking device 130 is used to implement this. The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and flash in the current frame. When not in tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in tracking state.
[0165] like Figure 6 As shown, the gaze-tracking camera captures left and right images of the user's left and right eyes. The captured images are then fed into a gaze-tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze-tracking system can continue capturing images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be fed into the pipeline for processing. However, in some embodiments or under certain conditions, not all captured frames are processed by the pipeline.
[0166] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, the image is analyzed to detect the user's pupil and flash, as indicated at 620. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0167] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flashes in part based on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flashes detected in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result is credible. For example, the result may be checked to determine whether a sufficient number of pupils and flashes used for gaze estimation were successfully tracked or detected in the current frame. At 650, if the result is not credible, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is credible, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.
[0168] Figure 6 This is intended as an example of an eye-tracking technology that can be used in a particular specific implementation. As will be recognized by those skilled in the art, in a computer system 101 for providing an XR experience to a user, other eye-tracking technologies that are currently available or will be developed in the future may be used to replace or in combination with the flash-assisted eye-tracking technology described herein, depending on the various implementations.
[0169] In some implementations, a portion of the captured real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0170] Therefore, this description describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and virtual objects. For example, a three-dimensional environment may optionally include a representation of a table existing in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of a computer system or passively displayed via a transparent or semi-transparent display of a computer system). As previously described, the three-dimensional environment may optionally be a mixed reality system, wherein the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system may optionally be able to selectively display portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system may optionally be able to display virtual objects in the three-dimensional environment to appear as if the virtual objects exist in the real world (e.g., the physical environment) by placing virtual objects in the three-dimensional environment at corresponding locations in the real world that have corresponding positions in the three-dimensional environment. For example, the computer system may optionally display a vase such that the vase appears as if a real vase were placed on top of a table in the physical environment. In some implementations, a corresponding location in the three-dimensional environment has a corresponding location in the physical environment. Therefore, when a computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., such as at or near a user's hand or at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object were at or near a physical object in the physical environment (e.g., the virtual object is displayed in the three-dimensional environment at a location in the physical environment that would be displayed if the virtual object were a real object at that specific location).
[0171] In some implementations, real-world objects that exist in the physical environment and are displayed in a 3D environment (e.g., and / or visible via a display generation component) can interact with virtual objects that exist only in the 3D environment. For example, the 3D environment may include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.
[0172] In a three-dimensional environment (e.g., a real environment, a virtual environment, or a hybrid environment including both real and virtual objects), an object is sometimes referred to as having depth or simulated depth, or as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension other than height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to a user's position or viewpoint, in which case the depth dimension varies based on the user's position and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the user's location relative to a surface of the environment (e.g., the surface of the environment's floor or ground), objects further away from the user along lines extending parallel to the surface are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder extending from the user's head toward the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., a direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), objects further away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from a line extending from and parallel to the user's viewpoint (e.g., defining depth in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere extending outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application displaying application and / or system content), where the user interface container has a height and / or width, and depth is a dimension orthogonal to the height and / or width of the user interface container. In some implementations, when a depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the container's depth dimension extends outward away from the user or the user's viewpoint), the container's height and / or width are typically orthogonal or substantially orthogonal to a straight line extending from the user's location (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some implementations, when a depth is defined relative to a user interface container, the object's depth relative to the user interface container refers to the object's positioning along the depth dimension of the user interface container. In some implementations, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending away from the user or the user's viewpoint in different directions and / or from different starting points).In some implementations, when depth is defined relative to a user interface container, the orientation of the depth dimension remains constant relative to the user interface container as the position of the user interface container changes, or as the user and / or the user's viewpoint changes (e.g., when multiple different viewers are viewing the same container in a 3D environment, such as during a collaborative session and / or when multiple participants are in a real-time communication session with shared virtual content including the container). In some implementations, for curved containers (e.g., containers including areas with curved surfaces or curved contents), the depth dimension may optionally extend into the surface of the curved container. In some cases, z-interval (e.g., the distance between two objects in the depth dimension), z-height (e.g., the distance of one object from another in the depth dimension), z-position (e.g., the position of an object in the depth dimension), z-depth (e.g., the position of an object in the depth dimension), or simulated z-dimensionality (e.g., depth used as a dimension of an object, a dimension of the environment, an orientation in space, and / or an orientation in simulated space) are used to refer to the concept of depth as described above.
[0173] In some implementations, a user may optionally be able to interact with virtual objects in a 3D environment using one or both hands as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system may optionally capture one or both of the user's hands and display a representation of the user's hands in the 3D environment (e.g., in a manner similar to displaying real-world objects in the 3D environment described above). Alternatively, in some implementations, the user's hands may be visible via the display generation component, through the ability to see the physical environment through the user interface, due to the transparency / semi-transparency of a portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / semi-transparent surface or onto the user's eyes or into the user's field of view. Thus, in some implementations, the user's hands are displayed at corresponding locations in the 3D environment and are treated as if they were objects in the 3D environment that could interact with virtual objects in the 3D environment as if these virtual objects were physical objects in the physical environment. In some implementations, the computer system may update the display of the user's hand representation in the 3D environment in conjunction with the movement of the user's hands in the physical environment.
[0174] In some embodiments described below, the computer system may optionally determine the “effective” distance between a physical object in the physical world and a virtual object in a three-dimensional environment, for example, to determine whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, holding, or within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object may optionally include one or more of the following: a finger pressing a virtual button, a user’s hand grasping a virtual vase, a user’s hand clasped together to pinch / hold the application’s user interface, and two fingers performing any other type of interaction described herein. For example, the computer system may optionally determine the distance between the user’s hand and the virtual object when determining whether and / or how the user is interacting with the virtual object. In some embodiments, the computer system determines the distance between the user’s hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, if a user's one or both hands are located at a specific location in the physical world, the computer system may optionally capture the one or both hands and display them at a specific corresponding location in a three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if it were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment may be compared with the location of a virtual object of interest in the three-dimensional environment to determine the distance between the user's one or both hands and the virtual object. In some embodiments, the computer system may optionally determine the distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between a user's one or both hands and a virtual object, the computer system may optionally determine the corresponding location of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if it were a physical object rather than a virtual object), and then determine the distance between the corresponding physical location and the user's one or both hands. In some embodiments, the same technique may optionally be used to determine the distance between any physical object and any virtual object. Therefore, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system may optionally perform any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or map the position of the virtual object to the physical environment.
[0175] In some implementations, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what the physical stylus held by the user is pointing at. For example, if the user's gaze is directed at a specific location in the physical environment, the computer system may optionally determine a corresponding location in the three-dimensional environment (e.g., a virtual location of the gaze), and if a virtual object is located at that corresponding virtual location, the computer system may optionally determine that the user's gaze is directed at that virtual object. Similarly, the computer system may optionally be able to determine the direction in which the stylus is pointing in the physical environment based on the orientation of the physical stylus. In some implementations, based on this determination, the computer system determines a corresponding virtual location in the three-dimensional environment corresponding to the location pointed at by the stylus in the physical environment, and optionally determines that the stylus is pointing at the corresponding virtual location in the three-dimensional environment.
[0176] Similarly, the embodiments described herein may refer to the location of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the location of the computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system serves as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to a corresponding location in the three-dimensional environment. For example, the location of the computer system would be its location in the physical environment (and its corresponding location in the three-dimensional environment) such that, if the user stands at that location facing the corresponding portion of the physical environment visible via the display generation component, the user will see from that location objects in the physical environment that are positioned, oriented, and / or sized (e.g., in an absolute sense and / or relative to each other) in the same way as objects displayed or visible in the three-dimensional environment by or via the display generation component of the computer system. Similarly, if the virtual objects displayed in a 3D environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same location as these virtual objects in the 3D environment, and physical objects in the physical environment having the same size and orientation as in the 3D environment), then the position of the computer system and / or the user is the position from which the user will see these virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the 3D environment by the display generation components of the computer system.
[0177] In this disclosure, various input methods are described in relation to interaction with a computer system. When an example is provided using one input device or method, and another example is provided using another input device or method, it should be understood that each example is compatible with and optionally utilizes the input device or method described relative to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When an example is provided using one output device or method, and another example is provided using another output device or method, it should be understood that each example is compatible with and optionally utilizes the output device or method described relative to the other example. Similarly, various methods are described in relation to interaction with a virtual or mixed reality environment via a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described relative to the other example. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.
[0178] User interface and related processes Now turn our attention to implementations of user interfaces (“UIs”) and associated processes that can be implemented in computer systems, such as portable multifunction devices or head-mounted displays, in communication with display generation components and one or more input devices.
[0179] Figures 7A to 7SExamples include: a three-dimensional environment visible via a display generation component (e.g., display generation component 7100a or display generation component 120) of a computer system (e.g., computer system 101), and interactions occurring within the three-dimensional environment caused by user input directed to the three-dimensional environment and / or input received from other computer systems and / or sensors. In some embodiments, input is directed to a virtual object within the three-dimensional environment by a user gaze detected in an area occupied by the virtual object or by a hand gesture performed at a location in the physical environment corresponding to the area of the virtual object. In some embodiments, input is directed to a virtual object within the three-dimensional environment by a hand gesture performed when the virtual object has input focus (e.g., when the virtual object has been selected by concurrently and / or previously detected gaze input, by concurrently or previously detected pointer input, and / or by concurrently and / or previously detected gesture input) (e.g., optionally, at a location in the physical environment independent of the area of the virtual object in the three-dimensional environment). In some embodiments, input is directed to a virtual object within the 3D environment by an input device that has positioned a focus selector object (e.g., a pointer object or selector object) at the location of the virtual object. In some embodiments, input is directed to a virtual object within the 3D environment via other components (e.g., voice and / or control buttons). In some embodiments, input is directed to a physical object or a representation of a virtual object corresponding to a physical object by user hand movements (e.g., whole hand movement, whole hand movement in a corresponding posture, movement of a portion of the user's hand relative to another portion of the hand, and / or relative movement between the hands) and / or manipulations relative to a physical object (e.g., touch, swipe, tap, open, toward movement, and / or relative movement). In some embodiments, the computer system displays changes in the 3D environment based on input from sensors (e.g., image sensors, temperature sensors, biometric sensors, motion sensors, and / or proximity sensors) and contextual conditions (e.g., location, time, and / or the presence of other people in the environment) (e.g., displaying additional virtual content, stopping the display of existing virtual content, and / or transitioning between different levels of immersion in the display of visual content). In some implementations, the computer system displays some changes in the 3D environment based on input from other computers used by other users sharing the computer-generated environment with the user of the computer system (e.g., in a shared computer-generated experience, in a shared virtual environment, and / or in a shared virtual or augmented reality environment in a communication session). These changes include displaying additional virtual content, stopping the display of existing virtual content, and / or transitioning between different levels of immersion in the display of visual content.In some implementations, the computer system displays changes in a three-dimensional environment based on input from sensors (e.g., changes in the movement, deformation, and / or visual characteristics of user interfaces, virtual surfaces, user interface objects, and / or virtual landscapes), the sensors detecting movement of other people and objects, as well as movement of the user that may not conform to the criteria of recognized gesture inputs that trigger associated operations of the computer system.
[0180] In some embodiments, the 3D environment visible via the display generation components described herein is a virtual 3D environment that includes virtual objects and content at different virtual locations within a 3D environment without a representation of a physical environment. In some embodiments, the 3D environment is a mixed reality environment that displays virtual objects at different virtual locations within a 3D environment constrained by one or more physical aspects of the physical environment (e.g., the location and orientation of walls, floors, surfaces, direction of gravity, time of day, and / or spatial relationships between physical objects). In some embodiments, the 3D environment is an augmented reality environment that includes a representation of a physical environment. In some embodiments, the representation of the physical environment includes corresponding representations of physical objects and surfaces at different locations within the 3D environment, such that spatial relationships between different physical objects and surfaces in the physical environment are reflected by spatial relationships between representations of physical objects and surfaces in the 3D environment. In some embodiments, when a virtual object is positioned relative to a representation of physical objects and surfaces in the 3D environment, the virtual object appears to have a corresponding spatial relationship with the physical objects and surfaces in the physical environment. In some implementations, the computer system transitions between displaying different types of environments based on user input and / or contextual conditions (e.g., transitioning between presenting computer-generated environments or experiences with different levels of immersion, and adjusting the relative salience of audio / visual sensory inputs from virtual content and representations from the physical environment).
[0181] In some embodiments, the display generation component includes a pass-through portion in which a representation of the physical environment is displayed or visible. In some embodiments, the pass-through portion of the display generation component is a transparent or translucent (e.g., see-through) portion of the display generation component that displays at least a portion of the physical environment surrounding the user or within the user's field of view (sometimes referred to as "optical pass-through"). For example, the pass-through portion is a portion of a head-mounted display or head-up display that is made translucent (e.g., less than 50%, 40%, 30%, 20%, 15%, 10%, or 5% opacity) or transparent, allowing the user to view the real world around them without removing the head-mounted display or moving away from the head-up display. In some embodiments, when displaying a virtual or mixed reality environment, the pass-through portion gradually transitions from translucent or transparent to completely opaque. In some embodiments, the pass-through portion of the display generation component displays a real-time feed (sometimes referred to as "optical pass-through") of images or video of at least a portion of the physical environment captured by one or more cameras (e.g., a rear-facing camera of a mobile device or associated with a head-mounted display, or other cameras that feed image data to a computer system). In some embodiments, the one or more cameras are pointed at a portion of the physical environment that is directly in front of the user's eyes (e.g., behind the display generating component relative to the user). In some embodiments, the one or more cameras are pointed at a portion of the physical environment that is not directly in front of the user's eyes (e.g., in a different physical environment, or to the side or behind the user).
[0182] In some embodiments, when virtual objects are displayed at locations corresponding to the positions of one or more physical objects in a physical environment (e.g., locations in a virtual reality, mixed reality, or augmented reality environment), at least some of the virtual objects are displayed to replace a portion of the camera's live view (e.g., a portion of the physical environment captured in the live view) (e.g., replacing its display). In some embodiments, at least some of the virtual objects and content are projected onto a physical surface or blank space in the physical environment and are visible through a transparent portion of the display generating component (e.g., visible as part of the camera view of the physical environment, or visible through a transparent or semi-transparent portion of the display generating component). In some embodiments, at least some of the virtual objects and content are displayed as portions covering the display and obscuring at least a portion of the view of the physical environment visible through the transparent or semi-transparent portion of the display generating component.
[0183] In some implementations, the display generation component displays different views of the 3D environment based on user input or movement that alters the virtual positioning of the viewpoint relative to the 3D environment, changing the currently displayed view of the 3D environment. In some implementations, when the 3D environment is virtual, the viewpoint moves based on navigation or motion requests (e.g., hand gestures in the air and / or gestures performed by moving one part of the hand relative to another part of the hand), without requiring movement of the user's head, torso, and / or the display generation component in the physical environment. In some implementations, movement of the user's head and / or torso, and / or movement of the display generation component or other position sensing elements of the computer system (e.g., due to the user holding the display generation component or wearing an HMD) relative to the physical environment causes a corresponding movement of the viewpoint relative to the 3D environment (e.g., with a corresponding change in direction, distance, speed, and / or orientation), resulting in a corresponding change in the currently displayed view of the 3D environment. In some embodiments, when a virtual object has a preset spatial relationship relative to a viewpoint (e.g., anchored or fixed to the viewpoint), movement of the viewpoint relative to the 3D environment will cause movement of the virtual object relative to the 3D environment while maintaining the virtual object's position within the field of view (e.g., the virtual object is referred to as head-locked). In some embodiments, the virtual object is body-locked to the user and moves relative to the 3D environment as the user moves as a whole in the physical environment (e.g., carrying or wearing display generation components and / or other position sensing components of the computer system), but will not move in the 3D environment in response to individual user head movements (e.g., rotation of display generation components and / or other position sensing components of the computer system around a fixed position of the user in the physical environment). In some embodiments, the virtual object may optionally be locked to another part of the user, such as the user's hand or wrist, and moves in the 3D environment according to movement of that part of the user in the physical environment to maintain a preset spatial relationship between the virtual object's position and the virtual position of that part of the user in the 3D environment. In some implementations, virtual objects are locked to a preset portion of the field of view provided by the display generation component and move in the three-dimensional environment according to the movement of the field of view, regardless of the movement of the user that does not cause a change in the field of view.
[0184] In some implementations, the view of the 3D environment sometimes does not include a representation of the user's hands, arms, and / or wrists. In some implementations, such as Figures 7A to 7SAs shown and described herein with reference to these figures, the view of the three-dimensional environment includes a representation of the user's hands, arms, and / or wrists. In some embodiments, the representation of the user's hands, arms, and / or wrists is included in the view of the three-dimensional environment as part of a representation of the physical environment provided via a display generation component. In some embodiments, these representations are not part of the representation of the physical environment and are captured separately (e.g., pointed at the user's hands, arms, and wrists by one or more cameras) and displayed in the three-dimensional environment independently of the currently displayed view of the three-dimensional environment. In some embodiments, these representations include camera images captured by one or more cameras of a computer system or stylized versions of the arms, wrists, and / or hands based on information captured by various sensors. In some embodiments, these representations replace a portion of the display of the representation of the physical environment, overlay that portion of the representation of the physical environment, or obscure that portion of the view of the representation of the physical environment. In some embodiments, when the display generation component does not provide a view of the physical environment and provides a completely virtual environment (e.g., no camera view and no transparent pass-through portion), a real-time visual representation (e.g., a stylized representation or segmented camera image) of one or both of the user's arms, wrists, and / or hands may optionally still be displayed in the virtual environment. In some embodiments, if no representation of the user's hands is provided in the view of the three-dimensional environment, the positioning corresponding to the user's hands may optionally be indicated in the three-dimensional environment, for example, by altering the appearance of the virtual content at the positioning in the three-dimensional environment corresponding to the position of the user's hands in the physical environment (e.g., by altering the translucency and / or simulating changes in reflectivity). In some embodiments, the representation of the user's hands or wrists is outside the currently displayed view of the three-dimensional environment, while the virtual positioning in the three-dimensional environment corresponding to the position of the user's hands or wrists is outside the current field of view provided via the display generation component; and in response to the virtual positioning corresponding to the position of the user's hands or wrists moving within the current field of view due to movement of the display generation component, the user's hands or wrists, the user's head, and / or the user as a whole, the representation of the user's hands or wrists becomes visible in the view of the three-dimensional environment.
[0185] Figures 7A to 7S Examples of processing inputs performed using different numbers of input manipulators are shown.
[0186] Figure 7A A physical environment 7000 is shown, including a user 7002 interacting with a computer system 101. The computer system 101 is worn on the head of the user 7002 and is typically positioned in front of the user 7002. Figure 7AIn this system, user 7002 freely interacts with computer system 101 using their left hand 7020 and right hand 7022. The physical environment 7000 includes physical objects 7014, physical walls 7004 and 7006, and a physical floor 7008. (The text abruptly ends here.) Figures 7B to 7S As shown in the example, the display generation component 7100a of the computer system 101 is a head-mounted display (HMD) worn on the head of the user 7002 (e.g., in...). Figures 7A to 7S The content shown as visible via the display generation component 7100a of the computer system 101 corresponds to the environmental view of the user 7002 when wearing the head-mounted display.
[0187] In some implementations, the head-mounted display (HMD) 7100a includes one or more displays that show a representation of a portion of the three-dimensional environment 7000' corresponding to the user's viewpoint. While HMDs typically include multiple displays, including a display for the right eye and a separate display for the left eye showing a slightly different image to generate a user interface with stereoscopic depth, Figures 7B to 7S The image displays a single image corresponding to the image of a single eye and uses additional annotations or descriptions to indicate depth information. In some embodiments, the HMD 7100a includes one or more sensors (e.g., one or more inward-facing and / or outward-facing image sensors 314), such as sensor 7101a, sensor 7101b, and / or sensor 7101c. Figure 7B The one or more sensors are used to detect the user's state, including tracking the user's face and / or eyes (e.g., using one or more inward-facing sensors 7101a and / or 7101b) and / or tracking the user's hands, torso, or other movements (e.g., using one or more outward-facing sensors 7101c). In some embodiments, the HMD 7100a includes one or more input devices optionally located on the housing of the HMD 7100a, such as one or more buttons, a touchpad, a touchscreen, a scroll wheel, a rotatable and pressable digital crown, or other input devices. In some embodiments, the input element is a mechanical input element; in some embodiments, the input element is a solid-state input element that responds to a press input based on detected pressure or intensity. For example, in Figures 7B to 7S The HMD 7100a includes one or more of buttons 701, buttons 702, and a digital crown 703 for providing input to the HMD 7100a. It should be understood that additional and / or alternative input devices may be included in the HMD 7100a.
[0188] In some embodiments, the display generating component of computer system 101 is a touchscreen held by user 7002. In some embodiments, the display generating component is a stand-alone display, projector, or another type of display. In some embodiments, the computer system communicates with one or more input devices, including cameras or other sensors and input devices that detect movement of the user's hand, movement of the user's whole body, and / or movement of the user's head in the physical environment. In some embodiments, the one or more input devices detect movement of the user's hand, face, and / or whole body and current posture, orientation, and positioning. For example, in some embodiments, when the user's hand 7020 (e.g., left hand) is within the field of view of the one or more sensors of HMD 7100a (e.g., within the user's field of view), a representation of the user's hand 7020' is displayed in the user interface on the display of HMD 7100a (e.g., as a pass-through representation and / or as a virtual representation of the user's hand 7020). In some embodiments, when a user's hand 7022 (e.g., right hand) is within the field of view of one or more sensors of the HMD 7100a (e.g., within the user's field of view), a representation of the user's hand 7022' is displayed in the user interface on the display of the HMD 7100a (e.g., as a pass-through representation and / or as a virtual representation of the user's hand 7022). In some embodiments, the user's hand 7020 and / or the user's hand 7022 is used to optionally combine gaze input to perform one or more gestures (e.g., one or more air gestures). In some embodiments, one or more gestures performed using the user's hand 7020 and / or 7022 include direct air gesture input based on the positioning of the representation of the user's hand 7020' and / or 7022' displayed in the user interface on the display of the HMD 7100a. For example, direct air gesture input is determined to be pointing to a user interface object displayed at a location that intersects with the displayed location of the representation of the user's hand 7020' and / or 7022' in the user interface. In some embodiments, one or more gestures performed using the user's hand 7020 and / or 7022 include indirect air gesture input, which is based on a virtual object displayed at a location corresponding to the location where the user's attention is currently detected (e.g., and / or optionally not based on the location of the representation of the user's hand 7020' and / or 7022' displayed within the user interface). For example, when user attention to a user interface object is detected (e.g., based on gaze or other indication of user attention), indirect air gestures, such as gaze and pinch (e.g., or other gestures performed using the user's hand), are performed relative to the user interface object.
[0189] In some embodiments, user input is detected via a touch-sensitive surface or touchscreen. In some embodiments, one or more input devices include an eye-tracking component that detects the location and movement of the user's gaze. In some embodiments, a display generation component and optionally one or more input devices, along with a computer system, are part of a head-mounted device that moves and rotates with the user's head in the physical environment and changes the user's viewpoint in a three-dimensional environment provided via the display generation component. In some embodiments, the display generation component is a heads-up display that does not move or rotate with the user's head or entire body, but optionally changes the user's viewpoint in a three-dimensional environment based on the movement of the user's head or body relative to the display generation component. In some embodiments, the display generation component (e.g., a touchscreen) may optionally be moved and rotated by the user's hand relative to the physical environment or relative to the user's head, and changes the user's viewpoint in a three-dimensional environment based on the movement of the display generation component relative to the user's head or face or relative to the physical environment.
[0190] In some embodiments, one or more portions of the view of the physical environment 7000 visible to the user 7002 via the display generation component 7100a are digital pass-through portions, which include representations of corresponding portions of the physical environment 7000 captured by one or more image sensors of the computer system 101. In some embodiments, one or more portions of the view of the physical environment 7000 visible to the user 7002 via the display generation component 7100a are optical pass-through portions, because the user 7002 can see one or more portions of the physical environment 7000 through one or more transparent or translucent portions of the display generation component 7100a.
[0191] In some implementations, the computer system is configured to accept input (such as that referenced herein) Figures 7A to 7S The input described herein is processed as analog touch input or other analog input. For example, for reference in this document... Figures 7A to 7S The described input, detected as an air gesture (e.g., involving detecting the user's gaze and / or hand position and / or movement) rather than other types of input (e.g., touch input, mouse input, keyboard input, etc.), optionally provides information about the air gesture input using an input processing framework for handling other types of input, such as describing the air gesture input using input events and data structures for other types of input. Using a framework (such as an application programming interface (API)) provided by a system that converts other types of user input (such as air gestures) into other types of analog input allows applications designed for other types of interaction (such as touch interaction or interaction using a mouse and keyboard) to receive and process other types of user input without requiring application redesign and redeployment.
[0192] When touch-based input is used to control software running on a computer system via a touch-sensitive surface, the touch input has both temporal and spatial aspects. The temporal aspect includes the phase of the touch input, indicating when the touch has just begun, whether the touch is moving or stationary, and when the touch ends (e.g., when the finger providing the touch input is lifted off the touch-sensitive surface). The spatial aspect of the touch input includes the location within the user interface and / or a set of one or more user interface areas or user interface windows in which the touch input occurs. In an example touch input processing framework, touch input detected as one or more touch input signals via a touch-sensitive surface is represented by one or more input events (e.g., touch events) describing the temporal and spatial aspects of the touch input. For example, a corresponding touch input via the touch-sensitive surface may optionally be represented by a sequence of input events (e.g., touch events), and in some embodiments, multiple concurrent touch inputs via the touch-sensitive surface are represented by a sequence of input events (e.g., touch events) (e.g., where each input event (e.g., touch event) describes a corresponding portion of multiple concurrent touch inputs) or by a sequence of corresponding input events (e.g., touch events) (e.g., where each input event (e.g., touch event) describes a different portion of a different touch input). Therefore, a corresponding input event (e.g., a touch event) describes a corresponding portion of a current or recent touch input or multiple concurrent touch inputs (in which case, information about multiple concurrent touch inputs may optionally be arranged into one or more lists). Example information in a corresponding input event (e.g., a touch event) describing a corresponding touch input on or near a touch-sensitive surface includes the following items or subsets or supersets thereof: • An input identifier that identifies the input (e.g., in some implementations, the input identifier is the same for all input events (e.g., touch events) associated with the same input). • Timestamps of input events (e.g., timestamps of the corresponding parts of the input); • The location of the input, such as its location relative to an input element (e.g., using two-dimensional coordinates or, in some cases, three-dimensional coordinates, where the third coordinate represents the distance relative to the input element and can optionally be used to describe the height of the hover input above the input element). • The target of the input (e.g., its identifier is one or more user interface areas and / or user interface windows to which the input is directed). • The duration of the input (e.g., since the initial detection of an input identifier with that input event). • Input movement direction and / or speed; and / or • Indicates the current stage value. Here is an example of a stage value: ○ Hover start (e.g., indicating the start of input that is close to but does not touch the touch-sensitive surface). ○ Hover still (e.g., indicating that a hover input has been detected but is not in contact with the touch-sensitive surface and there is no change in proximity or position). ○ Hover changes (e.g., indicating an update on the proximity or position of a hover input that has been detected but is not in contact with a touch-sensitive surface). ○ Hover ends (e.g., indicating that touch input near or touching the touch-sensitive surface is no longer detected, such as when the hover input has moved out of the detection range of the touch-sensitive surface). ○ Input start (e.g., indicating initial contact with a touch-sensitive surface, such as contact with a touch-sensitive surface via a previously hovering touch input; or indicating the start of non-contact input). ○ Input is stationary (e.g., indicating that a detected contact is in progress and has not yet moved along the touch-sensitive surface); or it indicates that a detected non-contact input is in progress and has not yet moved). ○ Input change (e.g., indicating an update to the location of a detected contact with a touch-sensitive surface; or indicating an update to the location of a detected non-contact input); ○ Input termination (e.g., indicating that contact with the touch-sensitive surface has ceased to be detected, such as by lifting off the touch-sensitive surface as part of an end-of-touch gesture, or that the input has otherwise ended; or indicating the termination of detected non-contact input); and / or ○ Input cancellation (e.g., indicating that the input has been determined to be unexpected or has been otherwise identified as input that should be ignored, and that an application that performs an operation based on the input should cancel the operation, such as by restoring to the state before the input was detected).
[0193] In some implementations, the information in the input event includes information identifying the target location (e.g., the coordinates of the target location, an identifier of the object corresponding to the target location, etc.). In some implementations, in addition to the input target, the information in the input event also includes information identifying the target location (e.g., coordinates). In some implementations, the information identifying the target location is included in or can be derived from the information identifying the input target.
[0194] In some implementations, input (such as those referenced herein) is provided in the form of other types of input events (e.g., touch events). Figures 7A to 7SThe described input events (user input via hand 7020 and / or hand 7022) cause the user input (e.g., even if performed as an air gesture instead of touch input) to be processed as other types of input (e.g., simulated touch input, simulated mouse and / or keyboard input, or other types of input). Therefore, in some embodiments, the corresponding input event includes information describing the touch input corresponding to or simulating the input event. For example, in some embodiments, the corresponding input event includes the following or a subset or superset thereof: a touch identifier identifying the touch input, a timestamp of the touch event, the position of the touch input relative to the touch-sensitive surface, the target of the touch input, the duration of the touch input, the direction and / or speed of the touch input, and / or a stage value indicating the current stage of the touch input, such as touch start, touch stop, touch change, touch end, and touch cancellation.
[0195] In some implementations, one or more gesture recognizers are used to process input, such as those referenced herein. Figures 7A to 7SThe described input, wherein the one or more gesture recognizers are configured to monitor the progress of the input and determine whether the input matches a definition of a corresponding gesture that the computer system is configured to recognize. In some embodiments, the gesture recognizer transitions between one or more states from a plurality of states including "possible," "active," "end," "cancel," "fail," and / or other states, based on gesture information received (e.g., as one or more input events or one or more touch events) or otherwise acquired regarding the corresponding input. For example, selecting an input (such as a tap gesture) causes the tap gesture recognizer to transition from the "possible" state to the "end" state (e.g., when the execution of the tap gesture (such as by performing an air pinch or air tap gesture) has been recognized). In another example, scroll input causes the scroll gesture recognizer to transition from a "possible" state to an "active" state (during scrolling, such as when a finger performing an air pinch moves laterally while maintaining contact during the air pinch, or when a finger performing an air tap moves laterally while maintaining pressure during the air tap), and then to an "end" state (when the scroll input has ended, such as when an air pinch is released by separating the fingers or when an air tap is released by lifting the fingers) (e.g., both "active" and "end" indicate that the gesture has been recognized). In some implementations, the gesture recognizer remains in the "possible" state during a hover phase event (e.g., when the user indicates readiness for interaction but has not yet performed a portion of the interactive input, which determines whether the input can be recognized as a specific gesture or a portion thereof, or whether the input cannot be recognized as a specific gesture). In some implementations, the gesture recognizer transitions to a "cancelled" state in response to a "hover cancel," "input cancel," or "touch cancel" event. In some implementations, a gesture recognizer configured to recognize specific gestures transitions to a "failure" state in response to certain input events that make it impossible for the input to be recognized as a specific gesture (e.g., input movement causes a long press gesture recognizer to fail; input lasting longer than a threshold amount of time causes a tap gesture recognizer to fail; input ending too quickly causes a long press gesture recognizer to fail; input ending without movement causes a scroll gesture recognizer to fail). In some implementations, the gesture recognizer is part of a framework provided by the system, such as an application programming interface (API), which optionally handles input recognition in addition to converting certain types of user input (such as air gestures) into other types of input (such as touch input or mouse input).
[0196] Figure 7B An example is illustrated of the three-dimensional environment (e.g., at least partially corresponding to) visible to user 7002 via display generation component 7100a (also referred to herein as "HMD7100a") of computer system 101. Figure 7AThe view of the physical environment (7000) in the game. Figure 7B The three-dimensional environment may optionally include representations of objects in a physical environment (such as physical environment 7000) (e.g., captured by one or more cameras of computer system 101). For example, in Figure 7B In this context, the three-dimensional environment includes a representation 7014' of physical object 7014, representations 7004' and 7006' (respectively) of physical walls 7004 and 7006, and a representation 7008' of physical floor 7008. Furthermore, the three-dimensional environment includes one or more computer-generated objects, also referred to as virtual objects, such as the application user interface 7010 and objects 7012 and 7024 within the user interface 7010 (e.g., they are not representations of physical objects in the physical environment 7000). In some embodiments, the application user interface 7010 corresponds to the user interface of a software application (e.g., an email application, a web browser, a messaging application, a map application, or other software application) executing on the computer system 101. Figure 7B It is also shown that the representation 7020' of the user's left hand 7020 and the representation 7022' of the user's right hand 7022 are visible in a three-dimensional environment via the HMD 7100a and are ready to provide input (e.g., in a ready state).
[0197] Figure 7C Examples are given from Figure 7B Example transformations performed. Figure 7C An example of user input is illustrated, which includes at least a first portion (e.g., including making two fingers touch) of a hand 7020 (e.g., a representation 7020' of the user's hand 7020 visible via HMD 7100a (e.g., within a viewport in a three-dimensional environment) performing an air pinch gesture (also referred to herein as a "first air pinch gesture") when the user 7002's attention is directed toward an object 7012 in the user interface 7010 (e.g., based on the user 7002's gaze toward the object). Based on the user 7002's gaze position when the air pinch gesture performed by hand 7020 is detected, object 7012 is determined to be the target of the user input. In response to the detection of the air pinch gesture performed by hand 7020, computer system 101 delivers an input event (e.g., "Event 1") to an application (or, in some embodiments, system software) associated with the target object 7012 (or more generally associated with the user interface 7010 including the target object 7012). In some implementations, the input event "Event 1" includes information identifying the location of the user 7002's gaze as the current location of the detected user input. Optionally, as... Figure 7CAs shown, object 7012 is visually emphasized in appearance (e.g., highlighted and / or other visual emphasis associated with object targeting) to indicate that object 7012 is the current input target. Figure 7C It is also shown that timer 7016 is started when an air pinching gesture is detected by hand 7020.
[0198] Figures 7D to 7E Examples are given from Figure 7C An example transformation was performed, which involved performing operations in a three-dimensional environment using a single input manipulator (e.g., hand 7020). Figure 7D In the middle, the hand 7020 is already maintaining the pinching gesture in the air while relative to Figure 7C The hand 7020 in the middle is positioned to move downwards and to the left by a threshold distance d. th In response, computer system 101 delivers a subsequent input event (e.g., "Event 1B") to the application associated with target object 7012 (e.g., when hand 7020 has moved a threshold distance d). th arrive Figure 7E (Following the positioning of the hand 7020 shown), the subsequent input event has information identifying the position corresponding to the hand 7020 as the input position. Therefore, in some embodiments, although the input position is initially determined based on the position in the three-dimensional environment pointed to by the user 7002's gaze when the first air pinch gesture is detected, the input position is subsequently determined based on the current position of the hand 7020 and / or the amount of movement of the hand 7020 since the first air pinch gesture was detected. In some embodiments, the amount of movement d of the hand 7020 is... th The movement value d corresponding to the input position acc In some implementations, the input movement value relative to the movement of the hand 7020 is accelerated (e.g., linearly scaled (e.g., using a multiplier) or non-linearly scaled (e.g., using a non-linear function based on the amount or speed of the input manipulator movement) (e.g., d acc >d th ).
[0199] exist Figure 7E In response to the indicated hand movement distance d th arrive Figure 7E The input event "Event 1B" of the position of the hand 7020 shown indicates that the computer system 101 (e.g., more specifically, the application associated with the user interface 7010) will move the object 7012 (e.g., away from the user interface 7010) in the user interface 7010. Figure 7D As shown and by Figure 7EThe dashed outline 7046 in the user interface 7010 indicates the previous positioning of the object 7012 (the movement amount d corresponding to the movement input position, and optionally expanded) of the object 7012. acc .therefore, Figures 7D to 7E An example of an operation performed in a three-dimensional environment using a single input manipulator (e.g., hand 7020).
[0200] Figure 7F Examples are given from Figure 7C The alternative transition is performed after timer 7016 has expired (this indicates that more than a threshold amount of time T has elapsed since the first air pinch gesture was detected by hand 7020). th1 (Also referred to herein as the “first threshold time amount”) the time at which the second air pinch gesture performed by hand 7022 is detected (e.g., via the representation 7022' of the user’s hand 7022 visible to the user via HMD 7100a (e.g., within the viewport of a three-dimensional environment)). User input by hand 7020 and user input by hand 7022 do not meet the concurrency criterion because the two user inputs are within each other’s threshold time amount T. th1 The input is not detected (e.g., even if the user input meets other requirements of the concurrency criterion, such as pointing to the same gaze location in object 7012 and / or at least partially overlapping in time). Because the two user inputs do not meet the concurrency criterion, computer system 101 delivers an input event (e.g., "Event 2") for the user input made by hand 7022. Input event "Event 2" is delivered to the application associated with object 7012 (the target of the user input made by hand 7022, as determined based on the gaze location of user 7002 when a second air pinch gesture made by hand 7022 is detected), and optionally includes information identifying the gaze location of user 7002 as the current location 7026 of the user input made by hand 7022. Therefore, in Figure 7F In the scene, Figure 7C At least one input event "Event 1" that identifies the user 7002's gaze position as the input position has been delivered to the application associated with the user interface 7010 for user input by hand 7020 (e.g., as referenced herein). Figure 7C (as described), and Figure 7F At least one input event “Event 2” that identifies the gaze position of user 7002 as the input position has been delivered to the application associated with the user interface 7010 for user input by hand 7022 (e.g., due to failure to meet concurrency criteria, user input by hand 7020 and user input by hand 7022 are processed as different inputs using different input manipulators, but simultaneously have the same input position).
[0201] Figure 7G Examples are given from Figure 7C Another alternative transition is performed before timer 7016 has expired (this indicates that less than a threshold amount of time T has elapsed since the first air pinch gesture was detected by hand 7020). th1 The second air pinch gesture executed by hand 7022 was detected at the time. Figure 7G This indicates position 7026 in the three-dimensional environment corresponding to the current position of hand 7020 in physical space, and position 7028 in the three-dimensional environment corresponding to the current position of hand 7022 in physical space, wherein positions 7026 and 7028 are different. However, in Figure 7G In the example, user input from hands 7020 and 7022 has not yet met the concurrency criterion because hands 7020 and 7022 have not yet moved (e.g., moved at least a threshold distance d relative to each other). th However, user input made by hands 7020 and 7022 may still meet the concurrency criteria. Therefore, computer system 101 delays the delivery of input events for user input made by hand 7022, optionally until user input made by hands 7020 and 7022 meets the concurrency criteria, or until user input made by hands 7020 and 7022 can no longer meet the concurrency criteria (e.g., due to the elapsed time of one or more threshold amounts, thus failing to meet the concurrency criteria). Figure 7G It is also shown that timer 7018 is started when a second air pinch gesture is detected by hand 7022.
[0202] Figure 7H Examples are given from Figure 7G An example transition is performed where hand 7022 completes a second air pinch gesture (e.g., by breaking the contact between fingers that were in contact with each other during the second air pinch gesture), thus ending the user input performed by hand 7022. Figure 7H In the example, user input by hand 7020 (e.g., it continues because hand 7020 maintains the first air pinch gesture) and user input by hand 7022 failed to meet the concurrency criterion because user input by hand 7022 ended, and hands 7020 and 7022 did not move relative to each other by at least a threshold movement amount d. th (For example, when hand 7020 remains stationary, the threshold movement amount d is moved by hand 7022) th When hand 7022 remains stationary, the threshold movement amount d is moved by hand 7020. th Or, the total movement threshold d of both hand 7020 and hand 7022. thIn response to the detection of the end of the second air pinch gesture performed by hand 7022, computer system 101 delivers an input event (e.g., "Event 2") for the end of user input performed by hand 7022 to the application associated with object 7012 (the target of user input performed by hand 7022) (e.g., after an initial delay in transmitting the input event corresponding to the second air pinch gesture performed by hand 7022, as referenced herein). Figure 7G (as described).
[0203] Figure 7H The input event "Event 2" may optionally include information identifying the location of the input. In some implementations, the location of the input is initially determined based on the location in the three-dimensional environment to which the user 7002's gaze is directed when the second air pinch gesture is detected, and subsequently determined based on the current position of the hand 7022 and / or the amount of movement of the hand 7022 since the second air pinch gesture was detected. Figure 7H In the middle, position 7026 corresponding to the position of hand 7020 and position 7028 corresponding to the position of hand 7022 are... Figure 7G The same applies because hand 7020 has not moved since the first air pinch gesture was initiated, and hand 7022 has not moved since the second air pinch gesture was initiated. User 7002's gaze has also not moved since the second air pinch gesture was initiated. Therefore, regardless of whether the input position is based on the position of user 7002's gaze or on the position and / or movement of hand 7022, Figure 7H The input position included in the input event "Event 2" for user input performed by hand 7022 coincides with the gaze position of user 7002, because neither user 7002's gaze nor hand 7022 has moved since the second air pinch gesture was initiated.
[0204] Figure 7H It is also shown that in some embodiments, an input event "Event 2" is delivered in response to the detection of the end of user input by hand 7022, regardless of whether timer 7018 has expired, and therefore regardless of whether a second threshold time T has elapsed since the detection of the second air pinch gesture by hand 7022. th2 Therefore, in Figure 7H In the scene, Figure 7C At least one input event "Event 1" that identifies the user 7002's gaze position as the input position has been delivered to the application associated with the user interface 7010 for user input by hand 7020 (e.g., as referenced herein). Figure 7C (as described), and Figure 7HAt least one input event “Event 2” that identifies the gaze position of user 7002 as the input position has been delivered to the application associated with the user interface 7010 for user input by hand 7022 (e.g., user input by hand 7020 and user input by hand 7022 are processed as different inputs using different input manipulators due to failure to meet concurrency criteria).
[0205] In response to Figure 7H The input event “Event 2”, which indicates the end of user input directed by hand 7022 to target object 7012, is used by computer system 101 (e.g., more specifically, an application associated with user interface 7010) to perform a selection operation relative to object 7012, for example, by displaying object 7012 with an appearance that indicates object 7012 has been selected (e.g., with highlighting, selection outline and / or other visual emphasis associated with object selection).
[0206] Figure 7I Examples are given from Figure 7G An alternative transition is performed in which hand 7022 maintains the second air pinch gesture, while hand 7020 maintains the first air pinch gesture, and timer 7018 has expired, indicating that a second threshold time T has elapsed since the second air pinch gesture performed by hand 7022 was detected. th2 .exist Figure 7I In the middle, position 7026 corresponding to the position of hand 7020 and position 7028 corresponding to the position of hand 7022 are... Figure 7G The same applies because hand 7020 has not moved since initiating the first air pinch gesture, and hand 7022 has not moved since initiating the second air pinch gesture. In some implementations, it is therefore considered that the user input from hand 7020 and the user input from hand 7022 fail to meet the concurrency criterion because the user input from hand 7020 and the user input from hand 7022 are maintained for a second threshold time T. th2 Hands 7020 and 7022 did not move relative to each other by at least the threshold movement amount d. th Therefore, in response to the second threshold time T th2 Upon expiration, computer system 101 will deliver an input event (e.g., "Event 2") for user input performed by hand 7022 to the application associated with object 7012 (the target of user input performed by hand 7022) (e.g., after an initial delay in transmitting the input event corresponding to a second air pinch gesture performed by hand 7022, as referenced herein). Figure 7G (As described). Similar to Figure 7H The input event "Event 2" in the middle Figure 7IThe input event "Event 2" may optionally include information identifying the location of the input. Figure 7I In the example, this position coincides with the position pointed to by the user 7002's gaze (e.g., even if the input position is based on the position and / or movement of the hand 7022, since the hand 7022 has not moved since the second air pinch gesture was initiated). Therefore, in Figure 7I In the scene, Figure 7C At least one input event "Event 1" that identifies the user 7002's gaze position as the input position has been delivered to the application associated with the user interface 7010 for user input by the hand 7020, and Figure 7I At least one input event “Event 2” that identifies the gaze position of user 7002 as the input position has been delivered to the application associated with the user interface 7010 for user input by hand 7022 (e.g., user input by hand 7020 and user input by hand 7022 are processed as different inputs using different input manipulators due to failure to meet concurrency criteria).
[0207] Figure 7J Examples are given from Figure 7G Another alternative transition is performed, wherein while hand 7020 maintains the first air pinch gesture and hand 7022 maintains the second air pinch gesture, hand 7020 moves downward and to the left relative to hand 7022 by at least a threshold movement amount d. th (For example, hand 7020 moves while hand 7022 remains stationary). This is because before timer 7016 expires and therefore within the threshold time T. th1 The threshold movement d of hand 7020 relative to hand 7022 was detected internally. th Therefore, user input from hand 7020 and user input from hand 7022 satisfy concurrency criteria (e.g., consistent with some implementations where a threshold movement amount must be detected before timer 7016 has expired). In some implementations, the concurrency criteria require that a threshold movement amount d of hand 7020 relative to hand 7022 be detected while hand 7020 maintains a first air pinch gesture and hand 7022 maintains a second air pinch gesture, and before timer 7018 has expired. th Optionally, it is not required to detect the threshold movement d before timer 7016 has expired. th .
[0208] Based on the user input from hand 7020 and user input from hand 7022 satisfying the concurrency criterion, computer system 101 transmits a set of associated input events. For example... Figure 7J As shown, computer system 101 transmission cancelled. Figure 7CThe input event "Event 2A" is "Event 1" (for example, "Event 2A" is an "input cancellation" event (e.g., an input event including an "input cancellation" phase value)). Furthermore, the computer system 101 transmits "Event 2B" which has information describing user input performed by hand 7020, and "Event 2C" which has information describing user input performed by hand 7022. Therefore, the initial input event describing the first part of the user input performed by hand 7020 ( Figure 7C Event 1) is cancelled, causing the user input performed by hand 7020 to be processed as part of a composite user input performed using multiple input manipulators (e.g., together with the user input performed by hand 7022) (e.g., while the user input performed by hand 7020 continues to be processed as a single input manipulator user input (as referenced herein)). Figure 7D , Figure 7P and Figure 7R (As described) in contrast. Optionally, "Event 2A" may be transmitted separately from "Event 2B" and "Event 2C", or the information in "Event 2A" may be included in "Event 2B" and / or "Event 2C" without transmitting "Event 2A" separately.
[0209] Because user 7002's gaze points to the same location when user input from hand 7020 is first detected, and when user input from hand 7022 is first detected (or, in some embodiments, because user 7002's gaze remains in the same location during the period when user input from hand 7020 is first detected), the user inputs from hands 7020 and 7022 have a common point of interaction determined based on the location of user 7002's gaze, and therefore object 7012 in user interface 7010 is the target of both input events "Event 2B" and "Event 2C". Therefore, both input events "Event 2B" and "Event 2C" are transmitted to the application associated with user interface 7010. However, because the user inputs from hands 7020 and 7022 have been identified as part of a complex multi-input manipulator user input, information about the relative positions of hands 7020 and 7022 is needed in addition to information about user 7002's gaze location. Therefore, the input event "Event 2B" for hand 7020 includes information identifying the location of user input made by hand 7020 as location 7026 in the three-dimensional environment corresponding to the current location of hand 7020 (e.g., and not necessarily coinciding with the gaze location of user 7002), and the input event "Event 2C" for hand 7022 includes information identifying the location of user input made by hand 7022 as location 7028 in the three-dimensional environment corresponding to the current location of hand 7022 (e.g., and not necessarily coinciding with the gaze location of user 7002), wherein location 7026 and location 7028 are different.
[0210] Figure 7K Examples are given from Figure 7J The transformation was performed, and it was shown that the hand 7020 was moved downwards and to the left by at least a threshold distance d. th (exist Figure 7J (As shown in the image) After that, the position of the user input made by the hand 7020 has been changed to position 7030 in the three-dimensional environment, which corresponds to the position of the hand 7020 after moving a distance d. th The current position is then determined by the user input from the still-unmoved hand 7022, which remains the position 7028 in the three-dimensional environment. Figure 7KAn example is illustrated where user input from hands 7020 and 7022 satisfies a concurrency criterion, and in response to the resulting input events "Event 2A," "Event 2B," and "Event 2C," computer system 101 (e.g., more specifically, an application associated with a user interface 7010 including a target object 7012) performs a sizing operation relative to object 7012 based on the relative movement between hands 7020 and 7022 (e.g., based on movement of hand 7020 away from hand 7022 while hand 7022 remains stationary, resulting in an increase in the distance between hands 7020 and 7022). Therefore, Figure 7K The object 7012 is shown as an increase in scale relative to its previous scale (as indicated by the dashed outline 7046). In some embodiments, the amount of size adjustment of the object 7012 is based on the amount of relative movement of the hands 7020 and 7022 (e.g., where a larger change in the distance between the hands 7020 and 7022 results in a larger (or alternatively smaller) size adjustment of the object 7012, and a smaller change in the distance between the hands 7020 and 7022 results in a smaller (or alternatively larger) size adjustment of the object 7012) and / or based on the direction of relative movement between the hands 7020 and 7022 (e.g., where movement that brings the hands 7020 and 7022 closer together decreases (or alternatively increases) the scale or size of the object 7012, and movement that further separates the hands 7020 and 7022 increases (or alternatively decreases) the scale or size of the object 7012).
[0211] In some implementations, a size adjustment operation is performed relative to a corresponding point in object 7012 that corresponds to the common interaction point between user input made by hand 7020 and user input made by hand 7022 (e.g., the gaze position of user 7002, which is the same when user input made by hand 7020 is detected as when user input made by hand 7022 is detected) (e.g., centered on the corresponding point).
[0212] In some implementations, because user input performed by hand 7020 and hand 7022 meets the concurrency criterion, Figure 7J Input position 7026 to Figure 7K The input position 7030 in the middle Figure 7J The movement of the input movement value d0 shown may optionally exclude the distance d relative to the hand 7020. th The input position for movement is expanded. In contrast, as referenced in this article... Figures 7D to 7E As described, Figure 7D The input movement value d acc Including the same distance d relative to the hand 7020th The input position for movement is expanded because Figures 7D to 7E The user input performed by the hand 7020 is a single input manipulator input. Therefore, Figure 7J The input movement value d0 is less than Figure 7D The input movement value d acc Accordingly, in some embodiments, the position and movement of hands 7020 and 7022 are mapped to corresponding positions in the three-dimensional environment (e.g., positions 7026 and 7028) (e.g., optionally, without input position scaling), such that movement of hands 7020 and 7022 toward each other without crossing (e.g., without contact) alters the mapped positions in the three-dimensional environment, where the mapped positions do not intersect. For example, if hand 7020 moves toward hand 7022 instead of... Figure 7J If the hand 7020 moves away from the hand 7022 as described above, then position 7026 will change to move toward position 7028; however, if the hand 7020 does not cross or pass through the hand 7022, then position 7026 will not cross or pass through position 7028.
[0213] Figure 7L Examples are given from Figure 7G Another alternative transformation was carried out, and it was similar to... Figure 7J The difference is: Figure 7L Instruct hand 7020 to move downwards and to the left a distance d a (For example, a non-expanded input movement value d1 corresponding to the movement of input position 7026) and the hand 7022 moves upward and to the right by a distance d. b (For example, the input movement amount value d2 corresponding to the optional unexpanded movement of input position 7028), thus satisfying the concurrency criterion because the total movement amount d of hand 7020 relative to hand 7022 is d. a +d b For at least the threshold movement amount d th (For example, with) Figure 7J When hand 7022 remains stationary, only hand 7020 moves by at least a threshold movement amount d. th (To create a contrast). Figure 7L The input events "Event 2A", "Event 2B", and "Event 2C" are similar to Figure 7J Those input events in (for example, the difference is that, with) Figure 7J In contrast, input events "Event 2B" and "Event 2C" indicate different input positions corresponding to different positions and movements of hand 7020 and hand 7022, respectively.
[0214] Figure 7M Examples are given from Figure 7L The transformation that has taken place is similar to that from Figures 7J to 7K The transformation. Figure 7M As shown, when hands 7020 and 7022 move away from each other by at least a threshold distance d... th Subsequently (e.g., in response to input events "Event 2A", "Event 2B", and "Event 2C"), a size adjustment operation is performed relative to the target object 7012 (e.g., its size is adjusted relative to the input position from...). Figure 7L The user input positions 7026 (operated by hand 7020) and 7028 (operated by hand 7022) are respectively located at... Figure 7M The changes in positions 7032 and 7034 are consistent, and therefore consistent with the total input movement value d1+d2. Therefore... Figure 7M The object 7012 is displayed at an increased scale relative to the previous scale of the object 7012 (as indicated by the dashed outline 7046).
[0215] In some implementation schemes, such as Figure 7M As shown in the example, because Figure 7L The user input of the hand 7020 and hand 7022 has the same characteristics as... Figure 7J The same common interaction points (e.g., based on the fact that the user 7002's gaze position is the same when user input is detected by hand 7020 as when user input is detected by hand 7022), so relative to Figure 7M In object 7012 and Figure 7K The same point (e.g., centered on that point) is used to perform the size adjustment operation, and therefore Figure 7M User interface response and Figure 7K The same applies to the movement of hands 7020 and 7022. Figure 7L Zhongyu Figure 7J Different (for example, even after only the hand 7020 moves) Figure 7K The positioning of hands 7020 and 7022 differs after hands 7020 and 7022 move in opposite directions. Figure 7M (Positioning of hands 7020 and 7022 in the model). Therefore, in Figure 7M In the diagram, the spatial relationship between the resized object 7012 and the dashed outline 7046 (e.g., representing the size and positioning of object 7012 before the resizing operation) and... Figure 7K The sized object 7012 and the dashed outline 7046 have the same spatial relationship.
[0216] Figure 7N Examples are given from Figure 7G Another alternative transformation was carried out, and it was similar to... Figure 7L The difference is that, Figure 7NInstructing hands 7020 and 7022 to rotate around each other to perform a rotation gesture (e.g., with...). Figure 7L (They move toward or away from each other to perform sizing adjustments and create contrast). Figure 7N In the middle, hand 7020 moves downward (e.g., optionally along a curved path) a distance d via hand 7022. a And hand 7022 moves upward (e.g., optionally along a curved path) a distance d past hand 7020. b Therefore, the concurrency criterion is met because the total movement d of hand 7020 relative to hand 7022 is sufficient. a +d b For at least the threshold movement amount d th (Optionally, the hand rotates around the other hand by at least a threshold movement d while the other hand remains stationary) th It will also meet the concurrency standard. Figure 7N The input events "Event 2A", "Event 2B", and "Event 2C" are similar to Figure 7L Those input events in (for example, the difference is that, with) Figure 7L In contrast, input events "Event 2B" and "Event 2C" indicate different input positions corresponding to different positions and movements of hand 7020 and hand 7022, respectively.
[0217] Figure 7O Examples are given from Figure 7N The transformation that has been carried out is similar to the transformation from Figure 7L to Figure 7M The difference lies in the transformation. Figure 7O This illustrates that hands 7020 and 7022 move relative to each other by at least a threshold distance d. th Then, a rotation operation is performed relative to the object 7012 based on the rotation of hands 7020 and 7022 relative to each other (e.g., its rotation relative to the input position from...). Figure 7N The user input positions 7026 (operated by hand 7020) and 7028 (operated by hand 7022) are respectively located at... Figure 7O The changes in positions 7036 and 7038 are consistent (e.g., a larger rotation of hands 7020 and 7022 causes a larger (or alternatively smaller) rotation of object 7012, and a smaller rotation of hands 7020 and 7022 causes a smaller (or alternatively larger) rotation of object 7012; and / or a clockwise rotation of hands 7020 and 7022 causes object 7012 to rotate clockwise (or alternatively counterclockwise), and a counterclockwise rotation of hands 7020 and 7022 causes object 7012 to rotate counterclockwise (or alternatively clockwise)). Therefore, Figure 7OThe object 7012 in the image is shown rotated relative to the previous orientation of the object 7012 (as indicated by the dashed outline 7046).
[0218] Figures 7P to 7Q Examples are given from Figure 7C Another alternative transformation involves performing different operations relative to different target objects in response to user input executed using different input manipulators and not meeting concurrency criteria. Figure 7P In the middle, hand 7020 has completed the first air pinch gesture (e.g., by breaking...). Figure 7C The contact between the fingers in contact with the user interface 7010 (e.g., because object 7012 in the user interface 7010 is the target of user input by the hand 7020) indicates the end of user input by the hand 7020. In some embodiments, the computer system 101 delivers the input event "Event 1-2" to the application associated with the user interface 7010 (e.g., because object 7012 in the user interface 7010 is the target of user input by the hand 7020) to indicate the end of user input by the hand 7020. In some embodiments, the input event "Event 1-2" is... Figure 7C Following the input event "Event 1", updated information about user input performed by hand 7020 is provided, including the fact that the user input performed by hand 7020 has ended (e.g., input event "Event 1-2" is an "End of Input" event (e.g., an input event including an "End of Input" phase value)). Input event "Event 1-2" may optionally include information identifying the location of the input, determined based on the current position of hand 7020 and / or the amount of movement of hand 7020 since the first air pinch gesture was detected; however, since hand 7020 has not moved since the first air pinch gesture was detected when user 7002's gaze was directed at object 7012, the location of the user input included in input event "Event 1-2" will be different from the location of the input itself. Figure 7C The same as the input event "Event 1". In response to input events "Event 1-2" indicating the end of user input directed by hand 7020 to target object 7012, computer system 101 (e.g., more specifically, an application associated with user interface 7010) performs a selection operation relative to object 7012 (e.g., as referenced herein). Figure 7H (As described). Therefore, user input from hand 7020 is processed as a single input manipulator user input, without considering user input from hand 7022.
[0219] In addition, Figure 7PIn the process, user 7002 has shifted their gaze from object 7012 to object 7024, and a second air pinch gesture performed by hand 7022 is detected when user 7002's gaze is directed at position 7040 within object 7024. The user input performed by hand 7022, including the second air pinch gesture, does not meet the concurrency criterion with the user input performed by hand 7020 because, according to user 7002's gaze moving from object 7012 to object 7024, the user input performed by hand 7022 points to a different target than the target pointed to by the user input performed by hand 7020. Therefore, computer system 101 delivers an input event (e.g., "Event 2-1") (optionally including position 7040 as the input position) to the application associated with the target object 7024. Figure 7P In the example, the application is the same as the one associated with user interface 7010 to which input events "Event 1-2" are delivered. Those skilled in the art will recognize that if object 7024 is instead associated with a different part of a user interface different from the application associated with user interface 7010, then input event "Event 2" will be delivered to the second application instead of the application associated with user interface 7010. Alternatively, as... Figure 7P As shown, object 7024 is visually emphasized in appearance (e.g., highlighted and / or other visual emphasis associated with object aiming) to indicate that object 7024 is (e.g., user input by hand 7022) the current input target.
[0220] Figure 7Q Examples are given from Figure 7P An example transition is performed where hand 7022 completes a second air pinch gesture (e.g., by breaking contact between the fingers), thus ending the user input performed by hand 7022. In response to detecting the end of the second air pinch gesture performed by hand 7022, computer system 101 delivers an input event (e.g., "Event 2-2") for the end of the user input performed by hand 7022; the input event "Event 2-2" is delivered to the application associated with target object 7024, or more generally to the application associated with user interface 7010 including target object 7024. Therefore, computer system 101 (e.g., more specifically, the application associated with user interface 7010) performs a selection operation relative to object 7024 (e.g., as referenced herein). Figure 7H (as described), and optionally perform a deselection operation relative to object 7012 (e.g., by removing in response to) Figure 7PThe input events “Event 1-2” are applied to the object selection associated with at least some or all of the visual emphasis. Therefore, user input made by hand 7022 is processed as a single input manipulator user input, without considering earlier user input made by hand 7020.
[0221] Figures 7R to 7S Examples are given from Figure 7C Another alternative transformation involves performing different but concurrent operations relative to different target objects in response to user input executed using different input manipulators and not meeting concurrency criteria. Figure 7R In the middle, hand 7020 has maintained the first air pinch gesture for at least the threshold time T. th1 (For example, from) Figure 7C Since the scenario, and hand 7022 has initiated a second air pinch gesture when user 7002's gaze is directed at object 7024 (optionally within at least a threshold time T). th1 (After death). While maintaining the initial pinched hand gesture in the air, hand 7020 moves downwards and to the left a distance d. L While maintaining the second air-together pinching gesture, hand 7022 moves downwards and to the right by a distance d. R .
[0222] Figure 7R This illustrates an input event (e.g.,) where computer system 101 delivers information about a portion of the user input made by hand 7020. Figure 7R "Event 1B" in the text is similar to Figure 7D The input event "Event 1B" in the text includes... Figure 7R The movement shown. For example, Figure 7R The input event "Event 1B" indicates input positions 7026 to 7042. Figure 7S The change is based on, for example, the change is based on, Figure 7R The illustrated position and movement of the hand 7020. In some embodiments, the amount of movement d of the hand 7020 is... L The movement value d corresponding to the input position L,acc And the input movement value can optionally be scaled relative to the movement of the hand 7020 (e.g., linearly scaled (e.g., using a multiplier) or non-linearly scaled (e.g., using a non-linear function based on the amount or speed of the input manipulator movement)) (e.g., d L,acc >d L ).
[0223] Similarly, computer system 101 delivers an input event (e.g.,) containing information about the initial portion of user input performed by hand 7022 (e.g., initiation of a second air pinch gesture). Figure 7R "Event 2" in the text is similar to Figure 7P The input event “Event 2-1” (e.g., in response to this information, object 7024 may optionally be visually emphasized in appearance, such as by using highlighting and / or other visual emphasis associated with object aiming, to indicate that object 7024 is the current input target). Furthermore, the computer system 101 delivers input events containing information about subsequent portions of the user input made by hand 7022 (e.g., Figure 7R The "Event 2D" in the text, this subsequent part includes Figure 7R The movement shown. For example, Figure 7R The input event "Event 2D" indicates input positions 7040 to 7044 ( Figure 7S The change is based on, for example, the change is based on, Figure 7R The position and movement of the illustrated hand 7022. In some embodiments, the amount of movement d of the hand 7022 is... R The movement value d corresponding to the input position R,acc And the input movement value can optionally be scaled relative to the movement of the hand 7022 (e.g., linearly scaled (e.g., using a multiplier) or non-linearly scaled (e.g., using a non-linear function based on the amount or speed of the input manipulator movement)) (e.g., d R,acc >d R ).
[0224] Figure 7S Examples are given from Figure 7R An example transformation is performed in which computer system 101 (e.g., more specifically, an application associated with user interface 7010) performs a movement operation relative to object 7024 based on input event "Event 2D" and concurrently performs a movement operation relative to object 7012 based on input event "Event 1B". Figure 7S In response to the indicated hand movement distance d L arrive Figure 7E The input event "Event 1B" is the position of the hand 7020 shown, and the object 7012 is in the user interface 7010 (e.g., away from the position of the hand 7020). Figure 7R As shown and as by Figure 7S The dashed outline 7046 in the user interface 7010 indicates the previous positioning of the object 7012 (the movement amount d corresponding to the movement input position, and optionally expanded) of the object 7012. L,acc In response to the instruction, the hand 7022 moves a distance d. R arrive Figure 7S The input event "Event 2D" is the position of the hand 7022 shown, and the object 7024 is in the user interface 7010 (e.g., away from the position of the hand 7022). Figure 7R As shown and as by Figure 7SThe dashed outline 7048 in the user interface 7010 indicates the previous positioning of the object 7012 (the movement amount d corresponding to, and optionally expanding, the movement input position). R,acc .
[0225] exist Figure 7S In this context, the operation performed relative to object 7012 is independent of the operation performed relative to object 7024 because user input from hand 7020 and user input from hand 7022 do not meet the concurrency criterion. Therefore, and because in some embodiments, the movement amount d of object 7012... L,acc The amount of movement d relative to the hand 7020 L The size was increased, while the movement amount d of object 7024 was increased. R,acc The amount of movement d relative to hand 7022 R It is enlarged so that even if hands 7020 and 7022 do not cross or pass each other (e.g., if hands 7020 and 7022 move toward each other instead of at...), Figures 7R to 7S (Moving in the direction shown), objects 7012 and 7024 may also intersect or pass each other (e.g., because the input position of user input by hands 7020 and 7022 is not as described herein). Figures 7J to 7K Mapping is performed as described.
[0226] although Figure 7S The diagram illustrates objects 7012 and 7024 within the same user interface 7010 of the same application, and input events (e.g., “Event 1B”) and input events (e.g., “Event 2” and “Event 2D”) delivered to the same application for user input made by hand 7020 and by hand 7022, respectively. However, in some embodiments, if objects 7012 and 7024 are in different user interfaces associated with different applications, the computer system 101 will similarly deliver different input events to different applications (e.g., deliver “Event 1B” to a first application and “Event 2” and “Event 2D” to a second application different from the first application), and concurrently perform movement operations relative to object 7012 (e.g., in the user interface of the second application different from the user interface of the first application) with movement operations performed relative to object 7024 (e.g., in the user interface of the second application different from the user interface of the first application). Therefore, user input from hand 7020 is processed as a single input manipulator user input, ignoring user input from hand 7022, and user input from hand 7022 is processed as a single input manipulator user input, ignoring user input from hand 7020.
[0227] The following text refers to... Figures 8A to 8D The described method 800 provides information about Figures 7A to 7G Additional description.
[0228] Figures 8A to 8D This is a flowchart of an exemplary method 800 for processing inputs executed using different numbers of input manipulators, according to some embodiments. In some embodiments, method 800 is implemented in a computer system (e.g., Figure 1A The computer system 101 in the display generation component (e.g., ...) executes the operation at the computer system 101 in the display generation component. Figure 1A , Figure 3 and Figure 4 The display generating component 120 (e.g., a head-up display, head-mounted device, display, touchscreen, projector, etc.) and one or more input devices (e.g., one or more cameras (e.g., color sensors, infrared sensors, structured light scanners, and / or other depth-sensing cameras) pointing downwards at the user's hand, forward from the user's head, and / or facing the user; eye-tracking devices; controllers held by the user and / or worn by the user; and / or other input hardware). In some embodiments, method 800 is stored in a non-transitory (or transient) computer-readable storage medium and is handled by one or more processors of a computer system (such as one or more processors 202 of computer system 101). Figure 1A The instructions executed by the control 110 in method 800 are controlled by the control. Some operations in method 800 may be combined, and / or the order of some operations may be changed.
[0229] When a view of the environment is visible via a display generation component (e.g., the environment is a two-dimensional or three-dimensional environment including one or more computer-generated portions and optionally one or more transparent portions), the computer system displays (802) a user interface including one or more user interface objects (e.g., user interface 7010 including object 7012 and object 7024). Figure 7B The computer system detects (804) one or more inputs via the one or more input devices (e.g., as referenced herein). Figures 7B to 7S The user input described is performed by hand 7020 and / or hand 7022.
[0230] Based on the determination that the one or more inputs include a first input performed using a first input manipulator (e.g., in a physical environment, such as a user's first hand, contact on a touch-sensitive surface, a controller, a wand, a mouse, or other input manipulator) and a second input performed using a second input manipulator different from the first input manipulator (e.g., in a physical environment, such as a user's second hand, contact on a touch-sensitive surface, a controller, a wand, a mouse, or other input manipulator), wherein the first input and the second input satisfy a concurrency criterion (and in some embodiments, in response to detecting the one or more inputs) (806), the computer system: provides (808) a first input event for the first input to a first application, and provides (810) a second input event for the second input to the first application. The first input event includes information identifying a target location. The second input event includes information identifying a target location (e.g., the same target location identified by the first input event). For example, as referenced herein... Figures 7J to 7K , Figures 7L to 7M and Figures 7N to 7O As described, user input performed by hand 7020 and user input performed by hand 7022 satisfy the concurrency criterion, and the computer system provides the application associated with user interface 7010 with an input event "Event 2B" containing information about user input performed by hand 7020 and an input event "Event 2C" containing information about user input performed by hand 7022.
[0231] In some implementations, the first input corresponds to a first location in the user interface, and the first input event also includes identifying the first location (e.g., location 7026). Figure 7J , Figure 7L and Figure 7N Information related to the second input. In some implementations, the second input corresponds to a second location in the user interface, and the second input event also includes identifying the second location (e.g., location 7028). Figure 7J , Figure 7L and Figure 7N Information regarding the first input and / or second input. In some embodiments, the first application includes instructions for rendering or displaying at least a portion of the user interface (e.g., a computer system uses the first application to render or display at least a portion of the user interface, such as corresponding user interface objects corresponding to both a target location and / or a first location and a second location). In some embodiments, upon determining that the detected first and second inputs meet a concurrency criterion, the first application is provided with a first input event and / or a second input event (e.g., automatically, without requiring intermediate user requests for information such as a first location corresponding to a first input manipulator and / or a second location corresponding to a second input manipulator, as described herein). Figure 7J , Figure 7Land Figure 7N middle).
[0232] In some implementations, the method includes: displaying a user interface comprising one or more user interface objects when a view of the environment is visible via a display generation component; detecting one or more inputs via the one or more input devices; and performing a first operation based on determining that the one or more inputs include a first input performed using a first input manipulator and a second input performed using a second input manipulator different from the first input manipulator, and that the first input and the second input satisfy a concurrency criterion: the first operation is performed based on the first input, the second input, and a target location.
[0233] In cases where multiple inputs are performed using different input manipulators and meet concurrency criteria (such as inputs performed simultaneously or close in time using both of the user's hands and / or otherwise more generally meet the criteria for two-handed gestures), providing the corresponding application with different input events that identify the same target location but are associated with different input locations on different input manipulators allows the receiving application to distinguish concurrent inputs from different input manipulators, understand the relative positions and relative movements of the input manipulators, and use this information in response to perform user interface operations, rather than simply knowing the positions and movements of the input manipulators that are independent of each other. This increases the flexibility of input processing and the range of supported inputs, thereby making interaction with computer systems more efficient (e.g., by reducing the amount of input required to perform an operation and / or enabling operations to be performed without displaying additional controls).
[0234] In some implementations, based on the determination that the one or more inputs include a corresponding input executed using a single input manipulator, without receiving another input that satisfies the concurrency criterion (e.g., without receiving any other input concurrent with the corresponding input, such as receiving a first input without receiving a second input or any other concurrent input executed using a second input manipulator, or receiving input provided using one hand instead of two hands) (812), the computer system performs a corresponding operation corresponding to the corresponding input executed using a single input manipulator (e.g., without performing another operation corresponding to another input manipulator). For example, as referenced herein... Figures 7D to 7EAs described, object 7012 moves within user interface 7010 because user input by hand 7020 is performed using a single input manipulator, without receiving another user input (e.g., by hand 7022) that meets the concurrency criterion of the user input performed by hand 7020. When input is performed using a single input manipulator (such as a hand), performing the corresponding operation based on information about the single input manipulator (e.g., without needing to consider information about multiple input manipulators) simplifies input processing and reduces the amount of time required to perform operations on a computer system.
[0235] In some implementations, based on the determination that the one or more inputs include a first input executed using a first input manipulator and a second input executed using a second input manipulator, and that the first and second inputs do not satisfy the concurrency criterion (814), the computer system: executes a first operation corresponding to the first input executed using the first input manipulator; and executes a second operation corresponding to the second input executed using the second input manipulator. The second operation is different from the first operation. In some implementations, the first operation is executed independently of the second input. In some implementations, the second operation is executed independently of the first input. In some cases, the first and second operations are executed concurrently (e.g., when the first and second inputs are executed concurrently and / or overlap in time), while in some cases, the first and second operations are not executed concurrently (e.g., when the first and second inputs are executed at different times and / or do not overlap in time). For example, as referenced herein... Figures 7P to 7Q As described, user input from hands 7020 and 7022 does not meet the concurrency criterion, therefore a sequential selection operation is performed relative to a first object 7012 (e.g., the one targeted by user input from hand 7020), and then relative to object 7024 (e.g., the one targeted by user input from hand 7022). In another example, as referenced herein... Figures 7R to 7S As described, user inputs performed by hands 7020 and 7022 do not meet concurrency criteria, therefore concurrent movement operations are performed relative to object 7012 (e.g., targeted by user inputs performed by hand 7020) and object 7024 (e.g., targeted by user inputs performed by hand 7022). In cases where multiple inputs are performed using different input manipulators and do not meet concurrency criteria (e.g., inputs performed using both hands but at different times, or inputs performed at the same time or close in time but not meeting other criteria for two-handed gestures), performing different operations corresponding to the different input manipulators (e.g., without needing to coordinate information about multiple input manipulators) simplifies input processing and reduces the amount of time required to perform operations on a computer system.
[0236] In some implementations, based on determining that the one or more inputs include a first input and a second input, wherein the first input and the second input satisfy a concurrency criterion, the computer system performs operation (816) based on the relative movement between the first input manipulator and the second input manipulator (e.g., movement of the first input manipulator relative to a stationary second input manipulator, movement of the second input manipulator relative to a stationary first input manipulator, or movement of both the first input manipulator and the second input manipulator, the movement of which changes the distance or spacing between the first input manipulator and the second input manipulator). For example, as referenced herein... Figures 7J to 7K , Figures 7L to 7M as well as Figures 7N to 7O The composite multi-input manipulator described in the text performs different operations on object 7012 based on the relative movement between hand 7020 and hand 7022. When multiple inputs are performed using different input manipulators (such as the user's two hands) and concurrency criteria are met, performing two-handed operations, at least partially controlled by the relative movement between the two input manipulators, increases the range of supported inputs and thus the range of executable operations, thereby making interaction with the computer system more efficient (e.g., by reducing the number of inputs required to perform the operation and / or enabling the operation to be performed without displaying additional controls).
[0237] In some embodiments, performing the operation includes (818): performing the operation with a first operational value based on determining that a relative movement of a first magnitude has occurred between the first input manipulator and the second input manipulator; and performing the operation with a second operational value different from the first operational value based on determining that a relative movement of a second magnitude has occurred between the first input manipulator and the second input manipulator, wherein the second magnitude of the relative movement is different from the first magnitude of the relative movement (e.g., as referenced herein). Figures 7J to 7K , Figures 7L to 7M as well as Figures 7N to 7O (As described in the context of complex multi-input manipulators for user input). When multiple inputs are performed using different input manipulators (such as the user's two hands) and concurrency criteria are met, performing two-handed manipulation based on the operational magnitude values of the relative amounts of movement between the different input manipulators enables faster and better, more intuitive manipulation of objects in the user interface.
[0238] In some implementations, the target location corresponds (820) to a corresponding user interface object in the one or more user interface objects (e.g., object 7012 in user interface 7010 is...). Figures 7J to 7K , Figures 7L to 7M and Figures 7N to 7OThe composite multi-input manipulator (the target of user input) enables operations on a specific user interface object located at a shared interaction point among multiple input manipulators, reducing the amount of time required to perform operations on a computer system and enabling operations with a wider range to be performed without displaying additional controls.
[0239] In some embodiments, performing this operation includes (822) rotating the corresponding user interface object (e.g., rotating by a first amount of rotation based on a first amount of relative movement between the first and second input manipulators, and rotating by a second amount of rotation based on a second amount of relative movement between the first and second input manipulators). For example, a counterclockwise movement of the first input manipulator relative to the second input manipulator (and / or the second input manipulator relative to the first input manipulator) about a corresponding axis in the physical environment causes the corresponding user interface object to rotate counterclockwise (or clockwise in some embodiments) about a corresponding axis in the visible environment, while a clockwise movement of the first input manipulator relative to the second input manipulator (and / or the second input manipulator relative to the first input manipulator) about a corresponding axis in the physical environment causes the corresponding user interface object to rotate clockwise (or counterclockwise in some embodiments) about a corresponding axis in the physical environment. For example, as referenced herein... Figures 7N to 7O As described, a rotation operation is performed on object 7012 (e.g., based on the magnitude and / or direction of the relative movement between hand 7020 and hand 7022). This enables the rotation of a specific user interface object in response to inputs that meet concurrent criteria, performed using multiple input manipulators, reducing the amount of time required to perform such operations on a computer system, and enabling the performance of operations with an increased range without displaying additional controls.
[0240] In some embodiments, performing this operation includes (824) changing the scaling of the corresponding user interface object (e.g., changing a first scaling change amount based on a first magnitude of relative movement between the first and second input manipulators, and changing a different second scaling change amount based on a second magnitude of relative movement between the first and second input manipulators). For example, moving the first and second input manipulators closer together will shrink the corresponding user interface object (e.g., reduce the scale of the corresponding user interface object) (or in some embodiments, enlarge it (e.g., increase the scale of the corresponding user interface object)), while moving the first and second input manipulators further apart will enlarge the corresponding user interface object (or in some embodiments, shrink it). For example, as referenced herein... Figures 7J to 7K as well as Figures 7L to 7MThe described resizing (e.g., rescaling or scaling) operation on object 7012 (e.g., based on the magnitude and / or direction of the relative movement between hand 7020 and hand 7022) reduces the amount of time required to perform such operations on a computer system, enabling resizing, rescaling, or scaling of a specific user interface object in response to input that meets concurrency criteria using multiple input manipulators, and allows for operations with an expanded range to be performed without displaying additional controls.
[0241] In some implementations, performing this operation includes (826) translating the corresponding user interface object (e.g., translating a first amount based on a first amount of relative movement between the first input manipulator and the second input manipulator, and translating a second amount based on a second amount of relative movement between the first input manipulator and the second input manipulator). For example, if in Figure 7L If hands 7020 and 7022 move in the same direction as each other (e.g., rather than moving away from or toward each other to perform a scaling or resizing operation), then a translation operation can optionally be performed in user interface 7010 (e.g., such as by shifting objects 7012 and 7024, as well as other elements displayed in user interface 7010, while maintaining the relative spatial relationships between the various elements displayed in user interface 7010). Making it possible to translate specific user interface objects in response to input that meets concurrency criteria, performed using multiple input manipulators, reduces the amount of time required to perform such operations on a computer system and enables operations with an increased range to be performed without displaying additional controls.
[0242] In some implementations, detecting the one or more inputs includes (828) detecting one or more air gestures (e.g., air pinch, air pinch and drag, or other air gestures). For example, this document references... Figures 7C to 7S The user inputs described for hands 7020 and 7022 include various air gestures, such as air pinch gestures and air pinch-and-drag gestures (e.g., hand movement while maintaining an air pinch gesture). In some embodiments, detecting one or more of the inputs includes detecting multiple (e.g., two or more) air gestures, such as two (or more) air pinch gestures (e.g., as referenced herein). Figure 7F (as described), air pinch gestures and air pinch and drag gestures (e.g., as referenced in this article). Figure 7J As described, two (or more) air pinch and drag gestures (e.g., as referenced in this article). Figure 7L , Figure 7N and Figure 7R(as described) or other combinations of air gestures. Operations are performed in response to one or more air gestures, and specifically in response to multiple air gestures that meet concurrency criteria, performed using different input manipulators (such as the user's two hands), increasing the range of supported inputs and thus enabling faster and more intuitive operation on computer systems.
[0243] In some implementations, determining that the first input and the second input meet the concurrency criterion includes (830): determining that the second input was detected within a first threshold time period since the first input was detected (e.g., in a context as referenced herein). Figure 7G Instead Figure 7F The described threshold time quantity T th1 (Inner). In some implementations, the concurrency criterion requires that the first input and the second input point to the same target location (e.g., a target location corresponding to a respective user interface object). In some implementations, if the concurrency criterion is not met based on the first input pointing to a different target location than the target location pointed to by the second input, then the input event for the first input is transmitted to a different target location than the input event for the second input (e.g., and in response, different corresponding operations associated with different target locations are performed). Requiring the detection of multiple inputs executed using different input manipulators within a threshold amount of time to satisfy the concurrency criterion for performing multi-handed (e.g., two-handed) operations eliminates ambiguity with (e.g., discrete) inputs intended to perform different operations by different input manipulators. This increases the range of supported inputs and thus the range of executable operations, thereby reducing the amount of time required to perform operations on a computer system.
[0244] In some implementations, determining that the first input and the second input meet the concurrency criterion further includes (832): determining within a second threshold time period (e.g., since the detection of the first input or the earlier of the first input and the second input) (e.g., within the time frame described herein). Figure 7J Instead Figure 7H The described threshold time quantity T th1 (within), or since the detection of the second input or the later of the first and second inputs (e.g., in the context of references herein). Figure 7I The described threshold time quantity T th2 The internal mechanism detects a relative movement of a threshold amount between the first and second input manipulators. In some implementations, the second threshold time amount (e.g., T) th2 ) is the first threshold time quantity (e.g., T) th1(e.g., elapses concurrently with a first threshold time amount). In some implementations, the second threshold time amount differs from the first threshold time amount (e.g., begins after the first threshold time amount and / or is less than or greater than the first threshold time amount). Requiring multiple inputs executed using different input manipulators to move relative to each other by a threshold movement amount within a threshold time amount to satisfy a concurrency criterion for performing multi-handed (e.g., two-handed) operations eliminates ambiguity with (e.g., stationary) inputs intended to perform different operations by different input manipulators. This increases the range of supported inputs and thus the range of executable operations, thereby reducing the amount of time required to perform the operations on a computer system.
[0245] In some implementations, the first input corresponds to a first location in the user interface that is different from the target location (e.g., the first location corresponds to a first input manipulator and is included in the first input event). Figure 7G The location 7026 for user input via hand 7020 is different. Figure 7C The target position is determined based on the user's gaze position (7002), and the second input corresponds to a second position in the user interface that is different from the target position (e.g., the second position corresponds to a second input manipulator and is included in the second input event). Figure 7G The location 7028 for user input via hand 7022 is different. Figure 7G The target location is determined based on the user's gaze position (7002). In some implementations, the input event is transmitted to the corresponding user interface object even if at least one of the first or second positions is outside the corresponding user interface object (e.g., Figure 7J , Figure 7L and Figure 7N (object 7012 in the context of the target object), because the initial target location is in the corresponding user interface object (e.g., based on...). Figure 7C and Figure 7G The gaze position of user 7002 (within object 7012). For different inputs executed using different input manipulators and meeting concurrency criteria, different input events are provided associated with different, spaced-out input positions and with a common target position. This allows the receiving application to distinguish concurrent inputs from different input manipulators, understand the relative positions and relative movements of the input manipulators, and use this information when performing user interface operations in response. This increases the flexibility of input processing and the range of supported inputs, and thus increases the range of executable operations, thereby reducing the amount of time required to perform operations on the computer system.
[0246] In some implementations, the target location corresponds (836) to the location pointed to by the user's attention (e.g., based on gaze) in the environment. For example, the location in object 7012 is based on... Figure 7C and Figure 7G The position of user 7002 in object 7012 and in the middle Figures 7J to 7K as well as Figures 7L to 7M The target location for the sizing operation performed in the system (e.g., the common interaction point of hands 7020 and 7022). Selecting the target location for input (including different inputs performed using different input manipulators and meeting concurrency criteria) based on the location the user is looking at and / or more generally the location where the user's attention is directed (e.g., when a corresponding input is detected) makes target selection for the corresponding operation more intuitive and precise, reducing user errors and the amount of time required to perform the operation on the computer system.
[0247] In some embodiments, detecting the one or more inputs includes (838) detecting a first input executed using a first input manipulator before detecting a second input executed using a second input manipulator. In some embodiments, in response to detecting a first input executed using a first input manipulator before detecting a second input executed using a second input manipulator: the computer system provides a first application with a third input event corresponding to the first input manipulator (e.g., as referenced herein). Figure 7C The described input event is "Event 1". In the case of input being performed using an input manipulator (such as a hand) (e.g., before a second input being performed using a second input manipulator (such as another hand) is detected), an input event is provided that represents and includes information about the input, enabling the receiving application to begin processing and responding to the input more quickly, thereby providing prompt feedback about the state of the computer system.
[0248] In some implementations, detecting the one or more inputs includes (840) detecting a second input performed using a second input manipulator (e.g., after detecting a first input performed using a first input manipulator). In some implementations, in response to detecting a second input performed using a second input manipulator: the computer system provides a fourth input event to the first application (e.g., as referenced herein). Figure 7J , Figure 7L and Figure 7NThe described input event, "Event 2A," includes an instruction for causing the first application to cancel or ignore the third input event (e.g., the fourth input event cancels the input of the third input event). In some embodiments, the fourth input event is transmitted to the first application along with the transmission of the first and second input events corresponding to two different input manipulator positions. In some embodiments, the instruction for causing the first application to ignore the third input event is included in the first and / or second input events (rather than being included separately in the fourth input event). In some embodiments, the instruction for causing the first application to ignore the third input event is provided in response to determining that the first and second inputs meet a concurrency criterion (e.g., being detected within a threshold amount of each other and / or including movement relative to a threshold amount of each other). After providing an earlier input event for the first input performed using the first input manipulator, canceling the earlier input event in response to detecting a second input performed using the second input manipulator (and optionally, in conjunction with providing input events for the first and second inputs that meet the concurrency criterion) allows the receiving application to process the input more accurately to determine an appropriate user interface response, which provides improved feedback on the state of the computer system.
[0249] In some implementations, in response to detecting a second input performed using a second input manipulator (e.g., the start of the second input or at least a portion of the second input), the computer system delays (842) providing an input event for the second input to the first application (e.g., until it is determined whether a concurrency criterion is met, or until the end of the second input is detected without meeting the concurrency criterion). For example, as referenced herein. Figure 7G As described, the transmission of input events for at least the initial portion of user input performed by hand 7022 is delayed. In some embodiments, delaying the provision of input events for a second input to a first application includes abandoning the provision of input events for the second input to the first application for a corresponding amount of time (e.g., until it is determined whether a concurrency criterion is met, or until the end of the second input is detected without meeting the concurrency criterion). After providing an earlier input event for the first input performed using the first input manipulator, delaying the provision of input events for the detected second input performed using the second input manipulator to the corresponding application allows the computer system time to interpret the second input (e.g., to determine whether the second input is unexpected, whether the second input, together with the first input, meets the concurrency criterion, and accordingly how to accurately represent the second input (and relatedly the first input) in the input event). This reduces user errors by reducing the chance of the receiving application producing an unexpected user interface response and makes interaction with the computer system more efficient.
[0250] In some implementations, after delaying the provision of an input event for the second input to the first application, the computer system detects (844) the end of the second input performed using the second input manipulator; and in response to detecting the end of the second input, the computer system provides an input event for the second input to the first application based on determining that a concurrency criterion is not met. For example, as referenced herein... Figure 7H As described, in response to detecting the end of user input performed by hand 7022 (e.g., termination of the unfolding of a second air pinch gesture), an input event "Event 2" is delivered to the application associated with user interface 7010. In some embodiments, the input event for the second input indicates the occurrence of the second input and is based on a target position in the corresponding user interface object (e.g., not based on the position corresponding to the second input manipulator, because the concurrency criterion is not met). After delaying the provision of the input event for the second input performed using the second input manipulator, the input event for the second input is provided to the corresponding application in response to detecting the end of the second input that does not meet the concurrency criterion, allowing the receiving application to begin processing and responding to the second input more quickly, thereby providing prompt feedback on the state of the computer system.
[0251] In some implementations, after a delay in providing an input event for the second input to the first application, the computer system provides (846) an input event for the second input to the first application based on a determination that a concurrency criterion is not met. In some implementations, an input event for the second input indicating the occurrence of the second input is provided to the first application even if the end of the second input is not detected (e.g., in cases where the concurrency criterion is not met even if the second input continues to be detected, such as due to the cessation of detection of the first input, or insufficient movement by the first input and / or the second input detected within a threshold time amount). For example, as referenced herein... Figure 7I As described, in the second threshold time quantity T th2 After the event has passed, the input event "Event 2" is delivered to the application associated with the user interface 7010 (e.g., even if user input by hand 7022 is in progress). After delaying the delivery of the input event for the second input performed using the second input manipulator, once it has been determined that the concurrency criterion is not met, the input event for the second input is provided to the corresponding application, enabling the receiving application to begin processing and responding to the second input more quickly, thereby providing prompt feedback on the status of the computer system.
[0252] In some implementations, the first application is configured to perform (848) a selection operation relative to a corresponding user interface object (e.g., a corresponding user interface object in one or more user interface objects; the corresponding user interface object corresponds to and / or includes the target location) in response to receiving an input event for a second input. For example, as referenced herein... Figure 7H As described, in response to an input event "Event 2" delivered to the application associated with the user interface 7010, a selection operation is performed relative to object 7012. This enables the receiving application to perform a selection operation relative to the target object in response to an input event for a second input that does not meet concurrency criteria, resulting in an intuitive user interface response while reducing the amount of time required to perform the operation on the computer system.
[0253] In some implementations, the determination of the one or more inputs includes a first input executed using a first input manipulator and a second input executed using a second input manipulator, wherein the first and second inputs satisfy a concurrency criterion: the position of the first input manipulator is mapped (850) to the position of the first input, and the position of the second input manipulator is mapped to the position of the second input, such that relative movement between the first and second input manipulators does not intersect (e.g., along one or more axes in the physical environment) moves the first and second inputs relative to each other, while the positions of the first and second inputs do not intersect (e.g., along one or more corresponding axes in the environment visible through the computer system). References Figures 7J to 7K Described as: At least for complex user input performed using multiple input manipulators, a method is used to map input manipulator positions to input positions transmitted as part of the input event in a manner that avoids input position crossing before the input manipulators have crossed (e.g., with reference to this document). Figures 7R to 7SThe described independent user inputs may have overlapping positions (creating a contrast). In some implementations, in response to detecting a first and a second input that meet concurrency criteria, the computer system performs an operation (e.g., rotation, scaling, translation, or other operation) relative to the corresponding user interface object based on the relative movement between the first and second input manipulators. In some implementations, when the first and second input manipulators are brought closer together without overlapping, the positions of the first and second inputs move closer together without overlapping accordingly. In some implementations, if the first and second input manipulators do not move past each other, the positions of the first and second inputs do not move past each other accordingly. When performing an operation in an environment including a user interface that is at least partially controlled by relative movement between two input manipulators, associating the positions of these two input manipulators in physical space with corresponding first and second input positions in the environment, such that the movement of the input manipulators toward each other in physical space without overlapping or crossing, and the movement of the corresponding first and second input positions in the environment without overlapping or crossing, reduces the chance of generating user interface responses that are inconsistent with physical reality (e.g., an image that flips as the input manipulators approach each other before they cross), which reduces errors and makes interaction with computer systems more efficient.
[0254] In some embodiments, based on determining that the one or more inputs include a corresponding input performed using a single input manipulator (e.g., based on the user performing the input using one hand instead of two hands) (852), moving the single input manipulator by a first input movement amount causes a user interface output with a first user interface output amount value (e.g., moving, scrolling, translating, resizing, rotating, scaling, or performing another user interface operation) (e.g., the position of the single input manipulator is mapped to the position of the corresponding input such that moving the single input manipulator causes the position of the corresponding input to move, scale linearly (e.g., using a multiplier), or scale non-linearly (e.g., using a non-linear function based on the amount or speed of the input manipulator movement) to cause the corresponding user interface output). In some embodiments, a first relative input movement amount between the first input manipulator and the second input manipulator causes a user interface output with a second user interface output amount value less than the first user interface output amount value (e.g., moving, scrolling, translating, resizing, rotating, scaling, or performing another user interface operation). For example, when a single input manipulator moves by a first input movement amount, the resulting user interface output has a first value greater than the user interface output value resulting from the first input manipulator moving by the first input movement amount relative to a stationary second input manipulator (e.g., consistent with the fact that the corresponding input movement amount is greater than the first input movement amount relative to the stationary second input movement amount). More generally, in some embodiments, a single input manipulator moving by a corresponding amount causes a corresponding user interface output with a value that is amplified or increased relative to the user interface output value caused by the same corresponding movement amount produced by the first and second input manipulators relative to each other. For example, as referenced herein... Figures 7D to 7E As described, the input movement value d acc Optionally, the corresponding amount of movement relative to the hand 7020 is increased, while in contrast, as referenced herein... Figures 7J to 7K As described, the change in input position from position 7026 to position 7030 optionally does not involve an amplification of the corresponding movement relative to hand 7020. Responding to input from a single input manipulator and producing a user interface output with a greater amount of movement relative to the single input manipulator, while responding to input from two input manipulators and producing a user interface output with a smaller amount of movement relative to the same specific movement of the two input manipulators relative to each other, increases the responsiveness of the user interface to a single input manipulator. This improves ergonomics by requiring less movement for a given amount of user interface output, while coordinating user interface responses to more complex inputs using two input manipulators to increase accuracy, thereby reducing errors and making interaction with the computer system more efficient.
[0255] For purposes of explanation, the foregoing description has been given by reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible based on the teachings above. The embodiments were chosen and described to best elucidate the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention with various modifications suitable for the contemplated particular purpose, as well as the various described embodiments.
[0256] As described above, one aspect of this technology involves collecting and using data from various sources to improve input processing during a user's XR experience. This disclosure envisions that, in some instances, this collected data may include personal information that uniquely identifies or can be used to contact or locate specific individuals. Such personal information may include demographic data, location-based data, telephone numbers, email addresses, Twitter IDs, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.
[0257] This disclosure recognizes that the use of such personal information data in the techniques of this invention can be used to benefit users. For example, personal information data can be used to improve input processing during a user's XR experience. Furthermore, this disclosure envisions other uses for personal information data that benefit users. For example, health and fitness data can be used to provide insights into a user's overall health status or can be used as positive feedback for individuals using the technology to pursue health goals.
[0258] This disclosure anticipates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with robust privacy policies and / or privacy measures. Specifically, such entities should implement and adhere to privacy policies and measures that are recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. Such policies should be easily accessible to users and should be updated as the collection and / or use of data changes. Personal information from users should be collected for legitimate and reasonable entity purposes and should not be shared or sold outside of these legitimate purposes. Furthermore, such collection / sharing should be conducted only after receiving informed consent from users. Additionally, such entities should consider taking any necessary steps to protect and safeguard the right to access such personal information data and ensure that other entities with access to such personal information data comply with the privacy policies and procedures of those other entities. Moreover, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices. Furthermore, policies and measures should be adapted to the specific types of personal information data collected and / or accessed, and to applicable laws and standards, including considerations of specific jurisdictions. For example, in the United States, the collection or acquisition of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while in other countries, health data may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy measures should be advocated for different types of personal data in each country.
[0259] Regardless of the foregoing, this disclosure also contemplates implementation schemes that allow users to selectively block the use or access to personal information data. That is, this disclosure envisions providing hardware and / or software components to prevent or block access to such personal information data. For example, with regard to input processing during an XR experience, the inventive technology can be configured to allow users to opt-in or opt-out at any time during or after registration for the service to participate in the collection of personal information data. In another example, certain data about a user (such as detailed information about the user's characteristics that provide user input (e.g., eyes, hands, fingers, wrists, head, and / or other features) and / or the location of such features) is not initially provided to the software receiving the user input unless the software requires and explicitly requests such data. Furthermore, this disclosure envisions providing notifications related to access to or use of personal information. For example, users may be notified when downloading an application that their personal information data will be accessed, and then reminded again just before the application accesses the personal information data.
[0260] Furthermore, the intent of this disclosure is that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once data is no longer needed, this risk can be minimized by restricting data collection and deleting data. Additionally, and where applicable, including in certain health-related applications, data deidentification can be used to protect user privacy. Deidentification can be facilitated, where appropriate, by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data among users), and / or other methods.
[0261] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it also contemplates that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, input processing during an XR experience can be performed by inferring preferences based on non-personal information data or a minimal measure of personal information (such as content requested by a user's associated device, other non-personal information available to the service, or publicly available information).
Claims
1. A method, the method comprising: At the computer system that communicates with the display generation components and one or more input devices: When the view of the environment is visible via the display generation component, a user interface including one or more user interface objects is displayed; One or more inputs are detected via the one or more input devices; as well as, The determination of the one or more inputs includes a first input executed using a first input manipulator and a second input executed using a second input manipulator different from the first input manipulator, wherein the first input and the second input satisfy a concurrency criterion: Provide a first input event to a first application in response to the first input, wherein the first input event includes information identifying the target location; as well as Provide the first application with a second input event in response to the second input, wherein the second input event includes information identifying the target location.
2. The method according to claim 1, wherein the method comprises: If, upon determining that the one or more inputs include a corresponding input executed using a single input manipulator, and no other input satisfying the concurrency criterion is received, a corresponding operation corresponding to the corresponding input executed using the single input manipulator is performed.
3. The method according to any one of claims 1 to 2, wherein the method comprises: Based on the determination that the one or more inputs include the first input executed using the first input manipulator and the second input executed using the second input manipulator, and that the first input and the second input do not satisfy the concurrency criterion: Perform a first operation corresponding to the first input performed using the first input manipulator; as well as Perform a second operation corresponding to the second input performed using the second input manipulator, wherein the second operation is different from the first operation.
4. The method according to any one of claims 1 to 3, wherein the method comprises: An operation is performed based on the relative movement between the first input manipulator and the second input manipulator, wherein the first input and the second input satisfy the concurrency criterion.
5. The method of claim 4, wherein performing the operation comprises: Based on determining that a relative movement of a first magnitude has occurred between the first input manipulator and the second input manipulator, the operation is performed with a first operational magnitude. as well as Based on the determination that a relative movement of a second magnitude occurs between the first input manipulator and the second input manipulator, wherein the second magnitude of the relative movement is different from the first magnitude of the relative movement, the operation is performed with a second magnitude that is different from the first operation magnitude.
6. The method according to any one of claims 4 to 5, wherein the target location corresponds to a corresponding user interface object among the one or more user interface objects.
7. The method of claim 6, wherein performing the operation includes rotating the corresponding user interface object.
8. The method according to any one of claims 6 to 7, wherein performing the operation includes changing the scaling ratio of the corresponding user interface object.
9. The method according to any one of claims 6 to 8, wherein performing the operation includes translating the corresponding user interface object.
10. The method of any one of claims 1 to 9, wherein detecting the one or more inputs includes detecting one or more air gestures.
11. The method of any of claims 1-10, wherein determining that the first input and the second input satisfy the concurrency criteria comprises: It is determined that the second input was detected within a first threshold time period since the first input was detected.
12. The method of claim 11, wherein determining that the first input and the second input satisfy the concurrency criteria further comprises: It is determined that a relative movement of a threshold amount between the first input manipulator and the second input manipulator is detected within a second threshold time period.
13. The method according to any one of claims 1 to 12, wherein the first input corresponds to a first position in the user interface that is different from the target position, and the second input corresponds to a second position in the user interface that is different from the target position.
14. The method according to any one of claims 1 to 13, wherein the target location corresponds to the location in which the user's attention is directed in the environment.
15. The method of any one of claims 1 to 14, wherein detecting the one or more inputs includes detecting the first input performed using the first input manipulator before detecting the second input performed using the second input manipulator, and the method comprises: In response to detecting the first input executed using the first input manipulator before detecting the second input executed using the second input manipulator: Provide the first application with a third input event corresponding to the first input manipulator.
16. The method of claim 15, wherein detecting the one or more inputs includes detecting the second input performed using the second input manipulator, and the method comprises: In response to detecting the second input performed using the second input manipulator: A fourth input event is provided to the first application, the fourth input event including instructions for the first application to cancel or ignore the third input event.
17. The method according to any one of claims 15 to 16, the method comprising: In response to detecting the second input performed using the second input manipulator, the input event for the second input is delayed and provided to the first application.
18. The method of claim 17, wherein the method comprises: After delaying the provision of the input event for the second input to the first application, the end of the second input performed using the second input manipulator is detected; as well as In response to the detection of the end of the second input, and based on the determination that the concurrency criterion is not met, an input event for the second input is provided to the first application.
19. The method of claim 17, wherein the method comprises: After delaying the provision of the input event for the second input to the first application, the input event for the second input is provided to the first application based on the determination that the concurrency criterion is not met.
20. The method of any one of claims 17 to 19, wherein the first application is configured to: perform a selection operation relative to a corresponding user interface object in response to receiving the input event for the second input.
21. The method according to any one of claims 1 to 20, wherein: The determination of the one or more inputs includes a first input executed using the first input manipulator and a second input executed using the second input manipulator, wherein the first input and the second input satisfy the concurrency criterion: The position of the first input manipulator is mapped to the position of the first input, and the position of the second input manipulator is mapped to the position of the second input, such that relative movement between the first input manipulator and the second input manipulator, without the first input manipulator and the second input manipulator intersecting, causes the first input and the second input to move relative to each other, while the positions of the first input and the second input do not intersect.
22. The method according to claim 21, wherein: Based on the determination that the one or more inputs include corresponding inputs executed using a single input manipulator, the single input manipulator moving a first input movement amount causes a user interface output having a first user interface output amount value, wherein: The first relative input movement between the first input manipulator and the second input manipulator causes a second user interface output with a value less than the first user interface output value.
23. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 22.
24. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; as well as A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 22.
25. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 1 to 22.
26. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs comprising instructions for: When the view of the environment is visible via the display generation component, a user interface including one or more user interface objects is displayed; detecting one or more inputs via the one or more input devices; as well as, The determination of the one or more inputs includes a first input executed using a first input manipulator and a second input executed using a second input manipulator different from the first input manipulator, wherein the first input and the second input satisfy a concurrency criterion: Provide a first input event to a first application in response to the first input, wherein the first input event includes information identifying the target location; as well as Provide the first application with a second input event in response to the second input, wherein the second input event includes information identifying the target location.
27. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; as well as A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: When the view of the environment is visible via the display generation component, a user interface including one or more user interface objects is displayed; One or more inputs are detected via the one or more input devices; as well as, The determination of the one or more inputs includes a first input executed using a first input manipulator and a second input executed using a second input manipulator different from the first input manipulator, wherein the first input and the second input satisfy a concurrency criterion: Provide a first input event to a first application in response to the first input, wherein the first input event includes information identifying the target location; as well as Provide the first application with a second input event in response to the second input, wherein the second input event includes information identifying the target location.
28. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: A component for displaying a user interface including one or more user interface objects, enabled when the view of the environment is visible via the display generation component; A component for detecting one or more inputs via the one or more input devices; as well as, The component enabled based on determining that the one or more inputs include a first input executed using a first input manipulator and a second input executed using a second input manipulator different from the first input manipulator, wherein the first input and the second input satisfy a concurrency criterion: Provide a first input event to a first application in response to the first input, wherein the first input event includes information identifying the target location; as well as Provide the first application with a second input event in response to the second input, wherein the second input event includes information identifying the target location.