Method for moving an object in a three dimensional environment
The computer system addresses inefficiencies in augmented and virtual reality interactions by using eye and hand-tracking technologies to reduce user inputs and enhance feedback, improving usability and battery life.
Patent Information
- Application Number
- JP2025168483
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-23
- Filing Date
- 2025-10-06
- Publication Date
- 2026-02-03
AI Technical Summary
Existing methods and interfaces for interacting with augmented and virtual reality environments are cumbersome, inefficient, and complex, leading to a significant cognitive burden on users and unnecessary energy consumption.
A computer system with improved methods and interfaces that reduce the number and type of user inputs by using algorithms to move and manipulate objects within a three-dimensional environment, including eye-tracking, hand-tracking, and touch-sensitive technologies, providing intuitive interaction and efficient feedback.
Enhances usability and reduces user errors by streamlining interactions, improving battery life through efficient input methods, and providing enhanced visual and tactile feedback.
Smart Images

Figure 2026016434000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 261,556, filed September 23, 2021, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0002] The present invention generally relates to computer systems having a display generation component and one or more input devices that present a graphical user interface, including but not limited to electronic devices that facilitate the movement of objects within a three-dimensional environment. [Background technology]
[0003] The development of computer systems for augmented reality has progressed significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects, including digital images, video, text, icons, and control elements such as buttons and other graphics.
[0004] However, methods and interfaces for interacting with environments (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which manipulating virtual objects is complex and error-prone create a significant cognitive burden for users and detract from the experience of the virtual / augmented reality environment. In addition, these methods are unnecessarily time-consuming, thereby wasting energy. This latter consideration is particularly important in battery-operated devices. Summary of the Invention
[0005] Therefore, there is a need for a computer system having improved methods and interfaces for providing users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or type of inputs from the user by helping the user understand the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.
[0006] The above-mentioned deficiencies and other problems associated with user interfaces for computer systems having display generating components and one or more input devices are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generating components, the output devices including one or more tactile output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI through stylus and / or finger contacts and gestures on the touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI or the user's body as captured by cameras and other movement sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through the interactions optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game playing, making phone calls, video conferencing, emailing, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note taking, and / or digital video playback, and executable instructions to perform those functions are optionally contained on a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.
[0007] There is a need for electronic devices with improved methods and interfaces for interacting with objects in a three-dimensional environment. Such methods and interfaces can complement or replace conventional methods for interacting with objects in a three-dimensional environment. Such methods and interfaces reduce the number, extent, and / or type of input from a user, creating a more efficient human-machine interface.
[0008] In some embodiments, the electronic device uses different algorithms to move an object within the three-dimensional environment based on the direction of such movement. In some embodiments, the electronic device modifies the size of an object within the three-dimensional environment as the distance between the object and the user's viewpoint changes. In some embodiments, the electronic device selectively resists movement of an object when the object contacts another object within the three-dimensional environment. In some embodiments, the electronic device selectively adds an object to another object within the three-dimensional environment based on whether the other object is a valid drop target for the object. In some embodiments, the electronic device facilitates the movement of multiple objects simultaneously within the three-dimensional environment. In some embodiments, the electronic device facilitates throwing of an object within the three-dimensional environment.
[0009] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention. [Brief explanation of the drawings]
[0010] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:
[0011] [Figure 1] FIG. 1 is a block diagram illustrating a computer system operating environment for providing a CGR experience, according to some embodiments.
[0012] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's CGR experience, according to some embodiments.
[0013] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a user with visual components of a CGR experience, according to some embodiments.
[0014] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.
[0015] [Figure 5]FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.
[0016] [Figure 6A] 1 is a flowchart illustrating a glint-assisted gaze tracking pipeline, according to some embodiments.
[0017] [Figure 6B] 1 illustrates an exemplary environment for an electronic device for providing a CGR experience, according to some embodiments.
[0018] [Figure 7A] 1 illustrates an example of an electronic device that utilizes different algorithms for moving an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 7B] 1 illustrates an example of an electronic device that utilizes different algorithms for moving an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 7C] 1 illustrates an example of an electronic device that utilizes different algorithms for moving an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 7D] 1 illustrates an example of an electronic device that utilizes different algorithms for moving an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 7E] 1 illustrates an example of an electronic device that utilizes different algorithms for moving an object in different directions within a three-dimensional environment, according to some embodiments.
[0019] [Figure 8A] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8B]1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8C] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8D] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8E] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8F] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8G] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8H] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8I] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8J] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. [Figure 8K] 1 is a flowchart illustrating a method of utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments.
[0020] [Figure 9A] 1 illustrates an example of an electronic device that dynamically resizes (or does not resize) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 9B] 1 illustrates an example of an electronic device that dynamically resizes (or does not resize) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 9C] 1 illustrates an example of an electronic device that dynamically resizes (or does not resize) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 9D] 1 illustrates an example of an electronic device that dynamically resizes (or does not resize) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 9E] 1 illustrates an example of an electronic device that dynamically resizes (or does not resize) virtual objects in a three-dimensional environment, according to some embodiments.
[0021] [Figure 10A] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10B] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10C] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10D] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10E] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10F] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10G]1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10H] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. [Figure 10I] 1 is a flowchart illustrating a method for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments.
[0022] [Figure 11A] 1 illustrates an example of an electronic device that selectively resists movement of an object in a three-dimensional environment, according to some embodiments. [Figure 11B] 1 illustrates an example of an electronic device that selectively resists movement of an object in a three-dimensional environment, according to some embodiments. [Figure 11C] 1 illustrates an example of an electronic device that selectively resists movement of an object in a three-dimensional environment, according to some embodiments. [Figure 11D] 1 illustrates an example of an electronic device that selectively resists movement of an object in a three-dimensional environment, according to some embodiments. [Figure 11E] 1 illustrates an example of an electronic device that selectively resists movement of an object in a three-dimensional environment, according to some embodiments.
[0023] [Figure 12A] 1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments. [Figure 12B] 1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments. [Figure 12C] 1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments. [Figure 12D]1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments. [Figure 12E] 1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments. [Figure 12F] 1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments. [Figure 12G] 1 is a flowchart illustrating a method for selectively resisting movement of an object in a three-dimensional environment, according to some embodiments.
[0024] [Figure 13A] 1 illustrates an example of an electronic device for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 13B] 1 illustrates an example of an electronic device for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 13C] 1 illustrates an example of an electronic device for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 13D] 1 illustrates an example of an electronic device for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments.
[0025] [Figure 14A] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14B] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14C] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14D] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14E] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14F] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14G] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments. [Figure 14H] 1 is a flowchart illustrating a method for selectively adding respective objects to an object in a three-dimensional environment, according to some embodiments.
[0026] [Figure 15A] 1 illustrates an example of an electronic device that facilitates moving and / or positioning multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 15B] 1 illustrates an example of an electronic device that facilitates moving and / or positioning multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 15C] 1 illustrates an example of an electronic device that facilitates moving and / or positioning multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 15D] 1 illustrates an example of an electronic device that facilitates moving and / or positioning multiple virtual objects within a three-dimensional environment, according to some embodiments.
[0027] [Figure 16A] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16B]1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16C] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16D] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16E] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16F] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16G] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16H] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16I] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments. [Figure 16J] 1 is a flowchart illustrating a method for facilitating movement and / or positioning of multiple virtual objects within a three-dimensional environment, according to some embodiments.
[0028] [Figure 17A] 1 illustrates an example of an electronic device that facilitates throwing virtual objects within a three-dimensional environment, according to some embodiments. [Figure 17B] 1 illustrates an example of an electronic device that facilitates throwing virtual objects within a three-dimensional environment, according to some embodiments. [Figure 17C]1 illustrates an example of an electronic device that facilitates throwing virtual objects within a three-dimensional environment, according to some embodiments. [Figure 17D] 1 illustrates an example of an electronic device that facilitates throwing virtual objects within a three-dimensional environment, according to some embodiments.
[0029] [Figure 18A] 1 is a flowchart illustrating a method for facilitating throwing a virtual object within a three-dimensional environment, according to some embodiments. [Figure 18B] 1 is a flowchart illustrating a method for facilitating throwing a virtual object within a three-dimensional environment, according to some embodiments. [Figure 18C] 1 is a flowchart illustrating a method for facilitating throwing a virtual object within a three-dimensional environment, according to some embodiments. [Figure 18D] 1 is a flowchart illustrating a method for facilitating throwing a virtual object within a three-dimensional environment, according to some embodiments. [Figure 18E] 1 is a flowchart illustrating a method for facilitating throwing a virtual object within a three-dimensional environment, according to some embodiments. [Figure 18F] 1 is a flowchart illustrating a method for facilitating throwing a virtual object within a three-dimensional environment, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0030] The present disclosure relates to a user interface that provides a computer-generated reality (CGR) experience to a user, according to some embodiments.
[0031] The systems, methods, and GUIs described herein facilitate electronic device interaction with objects in a three-dimensional environment and provide improved ways to manipulate the objects.
[0032] In some embodiments, the computer system displays a virtual environment in a three-dimensional environment. In some embodiments, the virtual environment is displayed via a far-field process or a near-field process based on the geometry (e.g., size and / or shape) of the three-dimensional environment (e.g., optionally simulating the real-world environment surrounding the device). In some embodiments, the far-field process includes introducing the virtual environment from a location farthest from the user's viewpoint and gradually expanding the virtual environment toward the user's viewpoint (e.g., using information about the location, position, distance, etc. of objects in the environment). In some embodiments, the near-field process includes introducing the virtual environment from a location farthest from the user's viewpoint and expanding outward from the initial location (e.g., expanding the size of the virtual environment relative to the display generation component) without considering the distance and / or position of objects in the environment.
[0033] In some embodiments, the computer system displays a virtual environment and / or atmospheric effects within the three-dimensional environment. In some embodiments, displaying the atmospheric effects includes displaying one or more lighting and / or particle effects in the three-dimensional environment. In some embodiments, in response to detecting a movement of the device, portions of the virtual environment are de-highlighted, optionally without reducing the atmospheric effects. In some embodiments, in response to detecting a rotation of the user's body (e.g., simultaneously with the rotation of the device), the virtual environment is moved to a new location within the three-dimensional environment, optionally aligned with the user's body.
[0034] In some embodiments, the computer system displays the virtual environment simultaneously with the user interface of the application. In some embodiments, the user interface of the application may be moved into the virtual environment and treated as a virtual object residing in the virtual environment. In some embodiments, the user interface is automatically resized when moved into the virtual environment based on the distance of the user interface when moved into the virtual environment. In some embodiments, while displaying both the virtual environment and the user interface, the user may request that the user interface be displayed as an immersive environment. In some embodiments, in response to a request to display the user interface as an immersive environment, the previously displayed virtual environment is replaced with the immersive environment of the user interface.
[0035] In some embodiments, the computer system displays a three-dimensional environment having one or more virtual objects. In some embodiments, in response to detecting a movement input in a first direction, the computer system moves the virtual object in a first output direction using a first movement algorithm. In some embodiments, in response to detecting a movement input in a second, different direction, the computer system moves the virtual object in a second output direction using a second, different movement algorithm. In some embodiments, during movement, as the first object moves proximate to the second object, the first object becomes aligned with the second object.
[0036] In some embodiments, the computer system displays a three-dimensional environment having one or more virtual objects. In some embodiments, in response to detecting movement input directed at the object, the computer system resizes the object if the object is moved toward or away from a user's viewpoint. In some embodiments, in response to detecting movement of the user's viewpoint, the computer system does not resize the object even if the distance between the viewpoint and the object changes.
[0037] In some embodiments, a computer system displays a three-dimensional environment having one or more virtual objects. In some embodiments, in response to movement of the individual virtual object in a respective direction that encompasses another virtual object, movement of the individual virtual object is resisted when the individual virtual object contacts the other virtual object. In some embodiments, movement of the individual virtual object is resisted because the other virtual object is a valid drop target for the individual virtual object. In some embodiments, the individual virtual object is moved through the other virtual object in a respective direction when movement of the individual virtual object through the other virtual object exceeds a respective magnitude threshold.
[0038] In some embodiments, a computer system displays a three-dimensional environment having one or more virtual objects. In some embodiments, in response to moving an individual virtual object to another virtual object that is a valid drop target for the individual virtual object, the individual object is added to the other virtual object. In some embodiments, in response to moving an individual virtual object to another virtual object that is an invalid drop target for the individual virtual object, the individual virtual object is not added to the other virtual object but is returned to the individual location from which the individual virtual object was originally moved. In some embodiments, in response to moving an individual virtual object to a individual location in an empty space in the virtual environment, the individual virtual object is added to a newly created virtual object at the individual location in the empty space.
[0039] In some embodiments, the computer system displays a three-dimensional environment having one or more virtual objects. In some embodiments, in response to movement input directed at the plurality of objects, the computer system moves the plurality of objects together within the three-dimensional environment. In some embodiments, in response to detecting an end of the movement input, the computer system arranges objects of the plurality of objects separately within the three-dimensional environment. In some embodiments, the plurality of objects are arranged in a stacked arrangement while being moved.
[0040] In some embodiments, the computer system displays a three-dimensional environment having one or more virtual objects. In some embodiments, in response to a throwing input, the computer system moves the first object toward a second object if the second object was targeted as part of the throwing input. In some embodiments, if the second object was not targeted as part of the throwing input, the computer system moves the first object within the three-dimensional environment according to the speed and / or direction of the throwing input. In some embodiments, targeting the second object is based on the user's gaze and / or the direction of the throwing input.
[0041] The processes described below enhance the usability of the device and streamline the user-device interface (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device) through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional controls that are displayed, performing an operation without requiring further user input when a set of conditions is met, improving privacy and / or security, and / or other techniques. These techniques also reduce power usage and improve the device's battery life by allowing the user to use the device more quickly and efficiently.
[0042] 1-6 provide a description of an exemplary computer system for providing a CGR experience to a user (as described below with respect to methods 800, 1000, 1200, 1400, and 1600, and / or 1800). In some embodiments, the CGR experience is provided to a user via an operating environment 100 that includes a computer system 101, as shown in FIG. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a touch sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some embodiments, one or more of input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with display generation component 120 (e.g., within a head-mounted or handheld device).
[0043] When describing a CGR experience, various terms are used to individually refer to several related, but distinct, environments that a user senses and / or with which the user can interact (e.g., using inputs detected by computer system 101 that cause the computer system generating the CGR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to computer system 101 generating the CGR experience). The following is a subset of these terms:
[0044] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.
[0045] Computer-Generated Reality (or Extended Reality (XR)): In contrast, a Computer-Generated Reality (CGR) environment refers to a wholly or partially mimicked environment that people sense and / or interact with via electronic systems. In a CGR, a subset of a person's body movements or representations thereof are tracked, and one or more properties of one or more virtual objects simulated within the CGR environment are adjusted accordingly to behave according to at least one law of physics. For example, a CGR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to the property(ies) of a virtual object(s) in a CGR environment may be made in response to a representation of a body movement (e.g., a voice command). A person may sense and / or interact with a CGR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatially expansive audio environment that provides the perception of a point sound source in 3D space. In another example, audio objects may enable audio transparency that selectively incorporates ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, a person may sense and / or interact with only audio objects.
[0046] Examples of CGR include virtual reality and mixed reality.
[0047] Virtual Reality: A virtual reality (VR) environment refers to an emulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in the VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.
[0048] Mixed Reality: A mixed reality (MR) environment refers to a mimetic environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On the virtual continuum, a mixed reality environment is anywhere between, but not including, a fully physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items from the physical environment or representations thereof). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.
[0049] Examples of mixed reality include augmented reality and augmented virtuality.
[0050] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person using the system perceives the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment that are representations of the physical environment. The system composites the images or videos with virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as "pass-through video," meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into a physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. Augmented reality environments also refer to mimic environments in which a representation of a physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, a system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of a physical environment may be distorted by graphically modifying (e.g., enlarging) portions thereof, thereby rendering the modified portions a non-photorealistic, altered version of the originally captured image. As a further example, a representation of a physical environment may be distorted by graphically removing or obscuring portions thereof.
[0051] Augmented Virtual: An augmented virtual (AV) environment refers to a mimicking environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images taken of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.
[0052] Perspective-Locked Virtual Object: A virtual object is perspective-locked when the computer system displays the virtual object in the same location and / or position within the user's perspective, even as the user's perspective shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's perspective is locked to the forward-facing orientation of the user's head (e.g., the user's perspective is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's perspective remains fixed even as the user's line of sight moves without moving the user's head. In embodiments in which the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's perspective is the augmented reality view being presented to the user on the display generating component of the computer system. For example, a perspective-locked virtual object that is displayed in the upper left corner of the user's perspective when the user's perspective is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's perspective when the user's perspective changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which a viewpoint-locked virtual object is displayed in a user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."
[0053] Environment-Locked Virtual Object: A virtual object is environment-locked (alternatively "world-locked") when a computer system displays the virtual object at a location and / or position within a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) locations and / or objects within a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint shifts, the locations and / or objects within the environment relative to the user's viewpoint change, resulting in the environment-locked virtual object appearing at a different location and / or position within the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered within the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes more left-leaning within the user's viewpoint (e.g., the position of the tree within the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear more left-leaning within the user's viewpoint. In other words, the location and / or position at which the environment-locked virtual object appears within the user's viewpoint depends on the position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine a position at which to display an environment-locked virtual object in the user's viewpoint. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object) or can be locked to a moving portion of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body that moves independent of the user's viewpoint, such as the user's hand, wrist, arm, or leg), so that the virtual object moves as the viewpoint or part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0054] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed-following behavior, which reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed-following behavior, the computer system intentionally delays the movement of the virtual object when it detects movement of the reference point that the virtual object is following (e.g., a part of the environment, the viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 and 300 cm from the viewpoint). For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits delayed-following behavior, the device ignores small amounts of movement of the reference point (e.g., ignores movement of the reference point that is less than a threshold amount of movement, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a “delayed following” threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object that maintains a substantially fixed position relative to a reference point includes a virtual object that is displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, or 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or in front / behind relative to the position of the reference point).
[0055] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be a light guide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, controller 110 is configured to manage and coordinate the user's CGR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., the physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within the housing (e.g., physical housing) of one or more of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the foregoing.
[0056] In some embodiments, display generation component 120 is configured to provide a CGR experience (e.g., at least a visual component of the CGR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with reference to FIG. 3. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0057] According to some embodiments, the display generation component 120 provides a CGR experience to the user while the user is virtually and / or physically present within the scene 105 .
[0058] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on their head, their hand, etc.). Thus, display generating component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, display generating component 120 surrounds the user's field of view. In some embodiments, display generating component 120 is a handheld device (e.g., a smartphone or tablet) configured to present CGR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generating component 120 is a CGR chamber, housing, or room configured to present CGR content without the user wearing or holding display generating component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with CGR content triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the CGR content responses are displayed via the HMD. Similarly, a user interface showing interactions with CGR content triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).
[0059] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features have not been shown for the sake of brevity so as not to obscure more pertinent aspects of the exemplary embodiments disclosed herein.
[0060] 2 is a block diagram of an example controller 110, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE®, THUNDERBOLT®, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE®, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0061] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0062] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or a subset thereof, including optional operating system 230 and CGR experience module 240:
[0063] Operating system 230 includes instructions for handling various basic system services and performing hardware-dependent tasks. In some embodiments, CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for respective groups of one or more users). To that end, in various embodiments, CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.
[0064] 1 , and optionally one or more of input devices 125, output devices 155, sensors 190, and / or peripheral devices 195. To that end, in various embodiments, data acquisition unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0065] In some embodiments, tracking unit 244 is configured to map scene 105 and track the position / location of at least display generating component 120 relative to scene 105 of FIG. 1 , and optionally relative to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 244 includes hand tracking unit 243 and / or eye tracking unit 245. In some embodiments, hand tracking unit 243 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generating component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 243 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 245 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)) or relative to CGR content displayed via display generation component 120. Eye tracking unit 245 is described in more detail below with respect to FIG. 5.
[0066] In some embodiments, coordination unit 246 is configured to manage and coordinate the CGR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, coordination unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0067] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0068] Although the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 may be located within separate computing devices.
[0069] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular embodiments, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 2 can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary depending on implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0070] 3 is a block diagram of an example of a display generation component 120, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that, for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0071] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0072] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to a waveguide display, such as a diffractive, reflective, polarized, holographic, etc. For example, the HMD 120 includes a single CGR display. In another example, the HMD 120 includes a CGR display for each eye of the user. In some embodiments, one or more CGR displays 312 are capable of presenting MR or VR content. In some embodiments, one or more CGR displays 312 are capable of presenting MR or VR content.
[0073] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would view if the HMD 120 were not present (and may be referred to as scene cameras). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.
[0074] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and a CGR presentation module 340:
[0075] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to a user via one or more CGR displays 312. To that end, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.
[0076] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0077] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. To that end, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0078] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed) based on the media content data. To that end, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0079] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0080] Although the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 may be located within separate computing devices.
[0081] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0082] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 243 (FIG. 2) to track the location / position of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to display generating components 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined relative to the user's hand). In some embodiments, hand tracking device 140 is part of display generating components 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generating components 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0083] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.
[0084] In some embodiments, image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to controller 110, which extracts high-level information from the map data. This high-level information is provided, typically via an application program interface (API), to an application running on the controller, which drives display generation component 120 accordingly. For example, a user can interact with software running on controller 110 by moving their hand 408 and changing the posture of their hand.
[0085] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define a set of orthogonal x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z-component measured by the image sensor. Alternatively, the hand tracking device 440 can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurements, based on single or multiple cameras or other types of sensors.
[0086] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps including the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.
[0087] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. Pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to the pose and / or gesture information.
[0088] In some embodiments, the gesture includes an air gesture, which is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to another of the user's hands, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of the user's body part (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of the user's body part).
[0089] In some embodiments, input gestures used in various examples and embodiments described herein include air gestures performed by the movement of a user's finger(s) relative to other finger(s) or part(s) of the user's hand to interact with a CGR or XR environment (e.g., a virtual or mixed reality environment), according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).
[0090] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides a computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., gaze) to a user interface element in combination with (e.g., simultaneous with) movement of the user's finger(s) and / or hand to perform pinch and / or tap input, as described in more detail below.
[0091] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly at a user interface object in response to performing an input gesture with the user's hand at a position corresponding to the user interface object's position in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly at a user interface object in response to detecting the user's attention (e.g., gaze) to the user interface object while performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in the three-dimensional environment. For example, for a direct input gesture, a user can direct the user's input at a user interface object by initiating the gesture at or near a position corresponding to the user interface object's displayed position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from an outer edge of the option or a central portion of the option). For indirect input gestures, a user can direct their input to a user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user initiates an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position that does not correspond to the displayed position of the user interface object).
[0092] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment, according to some embodiments. For example, pinch inputs and tap inputs, as described below, are performed as air gestures.
[0093] In some embodiments, the pinch input is part of an air gesture, including one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture is the movement of two or more fingers of a hand so that they touch each other, optionally including a short break (e.g., within 0-1 second) after the contact. A long pinch gesture that is an air gesture includes moving two or more fingers of a hand so that they touch each other for at least a threshold amount of time (e.g., at least 1 second) before detecting a break in the contact. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., when two or more fingers are touching), and the long pinch gesture continues until a break in the contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected immediately in succession (e.g., within a predetermined period of time) of each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaking contact between two or more fingers), and performs a second pinch input within a predetermined period of time (e.g., within 1 second or 2 seconds) after releasing the first pinch input.
[0094] In some embodiments, a pinch-and-drag gesture that is an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., followed by) a drag input that changes the position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers to touch each other and moves the same hand to the second position in the air with a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of a user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using a first hand of the user and a second pinch input performed using the other hand (e.g., a second of the user's hands) in conjunction with performing the pinch input using the first hand. In some embodiments, a movement between a user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).
[0095] In some embodiments, a tap input (e.g., directed toward a user interface element) performed as an air gesture includes movement(s) of a user's finger(s) toward the user interface element, movement of a user's hand toward a user interface element, optionally with the user's finger(s) extended toward the user interface element, a downward movement of a user's finger (e.g., mimicking a mouse click action or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture, moving the finger or hand away from the user's viewpoint and / or toward the object that is the target of the tap input followed by an end of the movement. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward the object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).
[0096] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the device determines that the user's attention is directed to the portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, and / or requiring the gaze to be directed to the portion of the three-dimensional environment, and if one of the additional conditions is not met, the device determines that the user's attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).
[0097] In some embodiments, detection of a ready configuration of a user or a portion of a user is detected by a computer system, and detection of a ready configuration of the hands is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs performed with the hands (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand geometry (e.g., a pre-pinch geometry with the thumb and one or more fingers extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap geometry with one or more fingers extended and the palm facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., above the user's waist, moved toward an area in front of the user below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface is responsive to attentional (e.g., gaze) input.
[0098] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively, may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 440, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, by dedicated circuitry within the housing of the hand tracking device 402, or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.
[0099] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404, according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing gray levels as depth increases. The controller 110 processes these depth values to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.
[0100] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, center of the palm, end of the hand where it connects to the wrist, etc.), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine hand gestures performed by the hand or the current state of the hand, according to some embodiments.
[0101] FIG. 5 illustrates an exemplary embodiment of eye tracking device 130 (FIG. 1). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 245 (FIG. 2) to track the position and movement of a user's gaze relative to scene 105 or relative to CGR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed on a wearable frame, the head-mounted device includes both components for generating CGR content for viewing by the user and components for tracking the user's gaze relative to the CGR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the CGR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in combination with head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generating components.
[0102] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0103] As shown in FIG. 5 , in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visual light to pass through. The gaze tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images, generates gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.
[0104] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, central visual location, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 have been determined, images captured by the eye tracking camera can be processed using glint-assisted methods to determine the user's current visual axis and viewpoint relative to the display.
[0105] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and a gaze tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, a projector, etc.) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or alternatively, may be directed at the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).
[0106] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0107] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the CGR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.
[0108] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be employed.
[0109] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the location and angle of the eye tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0110] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.
[0111] FIG. 6A illustrates a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted gaze tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.
[0112] As shown in FIG. 6A, an eye-tracking camera may capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system may continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0113] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.
[0114] At 640, proceeding from element 410, the current frame is analyzed to track pupils and glints based in part on previous information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes) and the pupil and glint information is passed to element 680 to estimate the user's gaze point.
[0115] 6A is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide a user with a CGR experience according to various embodiments.
[0116] FIG. 6B illustrates an exemplary environment for electronic device 101 for providing a CGR experience, according to some embodiments. In FIG. 6B, real-world environment 602 includes electronic device 101, user 608, and real-world objects (e.g., table 604). As shown in FIG. 6B, electronic device 101 is optionally tripod-mounted or otherwise secured to real-world environment 602 so that one or more hands of user 608 are free (e.g., user 608 is optionally not holding device 101 with one or more hands). As described above, device 101 optionally has one or more groups of sensors located on different sides of device 101. For example, device 101 optionally includes sensor group 612-1 and sensor group 612-2 located on the “rear” and “front” sides of device 101, respectively (e.g., capable of capturing information from each side of device 101). As used herein, the front side of the device 101 is the side that faces the user 608 and the back side of the device 101 is the side that faces away from the user 608 .
[0117] In some embodiments, sensor group 612-2 includes an eye tracking unit (e.g., eye tracking unit 245 described above with reference to FIG. 2) that includes one or more sensors for tracking the eyes and / or gaze of a user, and the eye tracking unit can "watch" user 608 and track the eye(s) of user 608 in the manner described above. In some embodiments, the eye tracking unit of device 101 can capture the movement, orientation, and / or gaze of the eyes of user 608 and process the movement, orientation, and / or gaze as input.
[0118] In some embodiments, sensor group 612-1 includes a hand tracking unit (e.g., hand tracking unit 243 described above with reference to FIG. 2) that can track one or more hands of user 608 held on the “back” side of device 101, as shown in FIG. 6B. In some embodiments, a hand tracking unit is optionally included in sensor group 612-2 so that user 608 can additionally or alternatively hold one or more hands on the “front” side of device 101 while device 101 tracks the position of the one or more hands. As described above, the hand tracking unit of device 101 can capture the movements, positions, and / or gestures of one or more hands of user 608 and process the movements, positions, and / or gestures as input.
[0119] In some embodiments, sensor group 612-1 optionally includes one or more sensors (e.g., image sensor 404 described above with reference to FIG. 4 ) configured to capture images of real-world environment 602, including table 604. As described above, device 101 can capture images of portions (e.g., part or all) of real-world environment 602 and present the captured portions of real-world environment 602 to the user via one or more display generation components of device 101 (e.g., a display of device 101 optionally located on a side of device 101 facing the user, opposite the side of device 101 facing the captured portions of real-world environment 602).
[0120] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with a CGR experience, e.g., a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0121] Accordingly, the description herein describes several embodiments of three-dimensional environments (e.g., CGR environments) that include representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table present in a physical environment that is captured and displayed within the three-dimensional environment (e.g., actively via a camera and display of the electronic device, or passively via a transparent or translucent display of the electronic device). As mentioned above, the three-dimensional environment is optionally a mixed reality system based on a physical environment, where the three-dimensional environment is captured by one or more sensors of the device and displayed via a display generation component. As a mixed reality system, the device can optionally display portions and / or objects of the physical environment such that each of the portions and / or objects of the physical environment appears to exist within the three-dimensional environment displayed by the electronic device. Similarly, the device can optionally display virtual objects in the three-dimensional environment such that the virtual objects appear to exist within the real world (e.g., the physical environment) by placing the virtual objects at respective locations within the three-dimensional environment that have corresponding locations in the real world. For example, the device optionally displays a vase in a manner that makes it appear as if the real vase were placed on a table in the physical environment. In some embodiments, each location in the three-dimensional environment has a corresponding location in the physical environment. Thus, when a device is described as displaying a virtual object at a location distinct from a physical object (e.g., at or near the location of a user's hand, or on or near a physical table, etc.), the device displays the virtual object at a particular location in the three-dimensional environment in a manner that makes it appear as if the virtual object were at or near the physical object in the physical world (e.g., the virtual object would be displayed in a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object was displayed if the virtual object were a real object at that particular location).
[0122] In some embodiments, real-world objects present in the physical environment that are displayed in the three-dimensional environment can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table in the physical environment and the vase is a virtual object.
[0123] Similarly, a user can optionally use one or more hands to interact with virtual objects in the three-dimensional environment as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the device optionally capture one or more of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, due to the transparency / translucency of the user interface, or the projection of the user interface onto a transparent / translucent surface, or the portions of the display generating components displaying the projection of the user interface to the user's eyes or field of view of the user's eyes, the user's hands are visible through the display generating components by the ability to see the physical environment through the user interface. Thus, in some embodiments, the user's hands are displayed at discrete locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, a user can move their hands to cause representations of their hands in the three-dimensional environment to move in coordination with the movement of the user's hands.
[0124] In some of the embodiments described below, the device can optionally determine an “effective” distance between a physical object in the physical world and a virtual object in the three-dimensional environment, for example, to determine whether a physical object is interacting with a virtual object (e.g., whether a hand is touching, grabbing, holding, etc., a virtual object, or whether it is within a threshold distance from the virtual object). For example, when determining whether and / or how a user is interacting with a virtual object, the device determines the distance between the user's hand and the virtual object. In some embodiments, the device determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of a target virtual object in the three-dimensional environment. For example, one or more of the user's hands are placed at specific positions in the physical world, which the device optionally captures and displays at specific corresponding positions in the three-dimensional environment (e.g., positions in the three-dimensional environment where the hand would be displayed if the hand were a virtual hand rather than a physical hand). The position of the hand in the three-dimensional environment is optionally compared to the position of the target virtual object in the three-dimensional environment to determine the distance between the user's one or more hands and the virtual object. In some embodiments, the device optionally determines the distance between a physical object and a virtual object by comparing positions in the physical world (e.g., as opposed to comparing positions in a three-dimensional environment). For example, when determining the distance between one or more of a user's hands and a virtual object, the device optionally determines the corresponding location in the physical world of the virtual object (e.g., the position where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and the user's one or more hands.In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the device optionally performs any of the techniques described above to map the location of the physical object to the three-dimensional environment and / or to map the location of the virtual object to the physical world.
[0125] In some embodiments, the same or similar techniques are used to determine where and what a user's gaze is directed at and / or where and what a physical stylus held by the user is directed at. For example, if a user's gaze is directed at a particular position in the physical environment, the device optionally determines a corresponding position in the three-dimensional environment, and if a virtual object is located at that corresponding virtual position, the device optionally determines that the user's gaze is directed at that virtual object. Similarly, the device can optionally determine where the physical stylus is pointing in the physical world based on the orientation of the physical stylus. In some embodiments, based on this determination, the device determines a corresponding virtual position in the three-dimensional environment that corresponds to the location in the physical world where the stylus is pointing, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.
[0126] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a device) and / or the location of a device within a three-dimensional environment. In some embodiments, a user of a device is holding, wearing, or otherwise located at or near the electronic device. Thus, in some embodiments, the location of the device is used as a proxy for the location of the user. In some embodiments, the location of the device and / or user within the physical environment corresponds to a distinct location within the three-dimensional environment. In some embodiments, the distinct location is a location from which a “camera” or “view” of the three-dimensional environment extends. For example, if a user were to stand at a location facing a distinct portion of the physical environment displayed by the display generating components, the location of the device would be a location within the physical environment (and its corresponding location within the three-dimensional environment) where the user would see objects within the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the objects are displayed by the display generating components of the device. Similarly, if the virtual objects displayed in the three-dimensional environment were physical objects in the physical environment (e.g., the virtual objects are located in the same physical environment location and have the same physical environment size and orientation as in the three-dimensional environment), the location of the device and / or user is the position at which the user would see the virtual objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other and to real-world objects) as displayed by the display generation components of the device.
[0127] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.
[0128] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions that perform a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed. User Interface and Related Processes
[0129] We now turn our attention to embodiments of user interfaces (“UIs”) and associated processes that may be executed in a computer system, such as a portable multifunction device or a head-mounted device, equipped with display generating components, one or more input devices, and (optionally) one or more cameras.
[0130] 7A-7E show examples of electronic devices that utilize different algorithms for moving objects in different directions within a three-dimensional environment, according to some embodiments.
[0131] 7A shows electronic device 101 displaying a three-dimensional environment 702 from the perspective of user 726 shown in an overhead view (e.g., facing the back wall of the physical environment in which device 101 is located) via a display generating component (e.g., display generating component 120 of FIG. 1 ). As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes a display generating component (e.g., a touchscreen) and multiple image sensors (e.g., image sensor 314 of FIG. 3 ). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101 could use to capture one or more images of a user or a part of a user (e.g., one or more of the user's hands) while the user interacts with electronic device 101. In some embodiments, the user interfaces shown and described below may also be realized on a head-mounted display that includes display generating components that display the user interface or three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).
[0132] 7A , device 101 captures one or more images of the physical environment around device 101 (e.g., operating environment 100), including one or more objects in the physical environment around device 101. In some embodiments, device 101 displays a representation of the physical environment in three-dimensional environment 702. For example, three-dimensional environment 702 includes coffee table representation 722a (corresponding to table 722b in the overhead view) that is optionally a representation of a physical coffee table in the physical environment, and three-dimensional environment 702 includes sofa representation 724a (corresponding to sofa 724b in the overhead view) that is optionally a representation of a physical sofa in the physical environment.
[0133] 7A , three-dimensional environment 702 also includes virtual objects 706a (corresponding to object 706b in the overhead view) and 708a (corresponding to object 708b in the overhead view). Virtual object 706a is optionally at a relatively small distance from the viewpoint of user 726, and virtual object 708a is optionally at a relatively large distance from the viewpoint of user 726. Virtual objects 706a and / or 708a are optionally one or more of an application's user interface (e.g., a messaging user interface, a content browsing user interface, etc.), a three-dimensional object (e.g., a virtual clock, a virtual ball, a virtual car, etc.), or any other element displayed by device 101 that is not included in device 101's physical environment.
[0134] In some embodiments, device 101 uses different algorithms to control the movement of an object in different directions within three-dimensional environment 702, for example, different algorithms for moving an object toward or away from the viewpoint of user 726, or different algorithms for moving an object vertically or horizontally within three-dimensional environment 702. In an embodiment, device 101 uses a sensor (e.g., sensor 314) to detect one or more of the user's hand 705b providing the movement input, the user's shoulder 705a corresponding to hand 705b (e.g., the right shoulder if the right hand is providing the movement input), or one or more absolute and / or relative positions (e.g., relative to each other and / or relative to the object being moved in three-dimensional environment 702) of an object toward which the movement input is directed. In some embodiments, the movement of the object is based on the detected quantities. Details on how the detected quantities are optionally utilized by device 101 to control the movement of the object are provided with reference to method 800.
[0135] 7A , hand 703a provides movement input directed toward object 708a, and hand 703b provides movement input directed toward object 706a. Hand 703a optionally provides input to move object 708a closer to the viewpoint of user 726, and hand 703b optionally provides input to move object 706a farther from the viewpoint of user 726. In some embodiments, such movement input includes moving the user's hand toward or away from the body of user 726 while the user's hand is in a pinch hand shape (e.g., while the tips of the thumb and index finger of the hand are touching). For example, from FIGS. 7A-7B , device 101 optionally detects hand 703a moving toward the body of user 726 while in the pinch hand shape, and device 101 optionally detects hand 703b moving away from the body of user 726 while in the pinch hand shape. 7A-7E, it should be understood that such hands and inputs need not be detected simultaneously by device 101. Rather, in some embodiments, device 101 responds independently to the hands and / or inputs shown and described in response to independently detecting such hands and / or inputs.
[0136] In response to the movement input detected in FIGS. 7A-7B, device 101 moves objects 706a and 708a in three-dimensional environment 702 accordingly, as shown in FIG. 7B. In some embodiments, for a given magnitude of user's hand movement, device 101 moves the target object more if the input is to move the object toward the user's viewpoint and moves the target object less if the input is to move the object away from the user's viewpoint. In some embodiments, this difference is to avoid moving objects far enough away from the user's viewpoint in three-dimensional environment 702 that the object is no longer reasonably interactable from the user's current viewpoint (e.g., too small within the user's field of view for reasonable interaction). In FIGS. 7A-7B, hands 703a and 703b optionally have the same magnitude of movement, but in different directions, as described above. In response to a given magnitude of movement of hand 703a toward the body of user 726, device 101 moves object 708a substantially the entire distance from its original location to the viewpoint of user 726, as shown in the overhead view of Figure 7B. In response to the same given magnitude of movement of hand 703b away from the body of user 726, device 101 moves object 708a away from the viewpoint of user 726 a distance less than the distance covered by object 706a, as shown in the overhead view of Figure 7B.
[0137] Additionally, in some embodiments, device 101 controls the size of objects included in three-dimensional environment 702 based on the object's distance from the user's 726 viewpoint to prevent objects from consuming a large portion of the user's 726 field of view from their current viewpoint. Thus, in some embodiments, objects are associated with an appropriate or optimal size for their current distance from the user's 726 viewpoint, and device 101 automatically resizes the objects to fit their appropriate or optimal size. However, in some embodiments, device 101 does not adjust the object's size until user input to move the object is detected. For example, in FIG. 7A , object 706a is optionally larger than its appropriate or optimal size for its current distance from the user's 726 viewpoint, but device 101 has not yet automatically scaled object 706a down to its appropriate or optimal size. In response to detecting input provided by hand 703b to move object 706a within three-dimensional environment 702, device 101 optionally reduces the size of object 706a within three-dimensional environment 702, as shown in the overhead view of Figure 7B. The reduced size of object 706a optionally corresponds to a current distance of object 706a from a viewpoint of user 726. Further details regarding controlling the size of an object based on the distance of the object from a viewpoint of a user are described with reference to the series of diagrams and method 900 of Figure 8.
[0138] In some embodiments, device 101 applies different amounts of noise reduction to moving objects depending on the object's distance from the viewpoint of user 726. For example, in FIG. 7C , device 101 displays a three-dimensional environment 702 including object 712a (corresponding to object 712b in the overhead view), object 716a (corresponding to object 716b in the overhead view), and object 714a (corresponding to object 714b in the overhead view). Objects 712a and 716a are, optionally, two-dimensional objects (e.g., similar to objects 706a and 708a), and object 714a is, optionally, a three-dimensional object (e.g., a cube, a three-dimensional model of a car, etc.). In FIG. 7C , object 712a is farther from the viewpoint of user 726 than object 716a. Hand 703e is optionally currently providing movement input to object 716a, and hand 703c is optionally currently providing movement input to object 712a. Hands 703e and 703c optionally have the same amount of noise at their respective positions (e.g., the same amount of shaking, trembling, or vibration of the hands, as reflected in double arrows 707c and 707e having the same length / size). Because device 101 optionally applies more noise reduction to the resulting movement of object 712a controlled by hand 703c than to the resulting movement of object 716a controlled by hand 703e (e.g., because object 712a is farther from the viewpoint of user 726 than object 716a), noise at hand 703c's position optionally results in less movement of object 712a than movement of object 716a resulting from noise at hand 703e's position. This difference in noise-reduced movement is optionally reflected in double arrow 709a having a smaller length / magnitude than double arrow 709b.
[0139] In some embodiments, in addition to, or alternatively to, utilizing different algorithms for moving objects toward and away from the user's viewpoint, device 101 utilizes different algorithms for moving objects horizontally and vertically within three-dimensional environment 702. For example, in FIG. 7C , hand 703d is providing an upward vertical movement input directed toward object 714a, and hand 703c is providing a rightward horizontal movement input directed toward object 712a. In some embodiments, in response to a given amount of hand movement, device 101 moves the object more within three-dimensional environment 702 when the movement input is a vertical movement input than when the movement input is a horizontal movement input. In some embodiments, this difference is to reduce the strain a user may feel when moving their hands, because vertical hand movement may be more difficult than horizontal hand movement (e.g., due to gravity and / or anatomical reasons). For example, in FIG. 7C , the movement amounts of hands 703c and 703d are optionally the same. In response, as shown in FIG. 7D, device 101 has moved object 714a vertically in three-dimensional environment 702 more than it has moved object 712a horizontally in three-dimensional environment 702.
[0140] In some embodiments, when objects are being moved within the three-dimensional environment 702 (e.g., in response to indirect manual input that optionally occurs while the hand is beyond a threshold distance (e.g., 1, 3, 6, 12, 24, 36, 48, 60, or 72 cm) from the object it is controlling, as described in more detail with reference to method 800), their orientation is optionally controlled differently depending on whether the object being moved is a two-dimensional object or a three-dimensional object. For example, while a two-dimensional object is being moved within the three-dimensional environment, device 101 optionally adjusts the orientation of the object to remain perpendicular to the viewpoint of user 726. For example, in FIGS. 7C-7D , while object 712a is being moved horizontally, device 101 automatically adjusts the orientation of object 712a to remain perpendicular to the viewpoint of user 726 (e.g., as shown in the overhead view of the three-dimensional environment). In some embodiments, the normality of the orientation of object 712a relative to the viewpoint of user 726 is optionally maintained for movement in any direction within three-dimensional environment 702. In contrast, while the three-dimensional object is being moved within the three-dimensional environment, device 101 optionally maintains the relative orientation of a particular surface of the object with respect to objects or surfaces within three-dimensional environment 702. For example, in FIGS. 7C-7D , while object 714a is being moved vertically, device 101 controls the orientation of object 714a such that the bottom surface of object 714a remains parallel to the floor within the physical environment and / or three-dimensional environment 702. In some embodiments, the parallelism between the bottom surface of object 714a and the floor within the physical environment and / or three-dimensional environment 702 is optionally maintained for movement in any direction within three-dimensional environment 702.
[0141] In some embodiments, as described in more detail with reference to method 800, an object being moved via direct movement input, optionally occurring while the hand is less than a threshold distance (e.g., 1, 3, 6, 12, 24, 36, 48, 60, or 72 cm) from an object it is controlling, optionally rotates (e.g., pitch, yaw, and / or roll) freely according to corresponding rotational input provided by the hand (e.g., rotation of the hand about one or more axes during the movement input). This free rotation of the object, optionally, contrasts with the controlled orientation of the object described above with reference to indirect movement input. For example, in FIG. 7C , hand 703e is optionally providing direct movement input to object 716a. From FIG. 7C to FIG. 7D , hand 703e is optionally moving leftward, moving object 716a leftward within three-dimensional environment 702, and also providing rotational (e.g., pitch, yaw, and / or roll) input to change the orientation of object 716a, as shown in FIG. 7D . As shown in FIG. 7D, object 716a has rotated according to the rotation input provided by hand 703e, and object 716a does not remain perpendicular to the user's 726 viewpoint.
[0142] However, in some embodiments, upon detecting the end of the direct movement input (e.g., upon detecting the release of a pinch hand shape being made by the hand or upon detecting that the user's hand has moved more than a threshold distance from the controlled object), device 101 automatically adjusts the orientation of the controlled object according to the above-described rules that apply to the type of object (e.g., two-dimensional or three-dimensional). For example, in FIG. 7E, hand 703e drops object 716a into open space within three-dimensional environment 702. In response, device 101 automatically adjusts the orientation of object 716a to be perpendicular to the viewpoint of user 726 (e.g., different from the orientation of the object at the moment of dropping), as shown in FIG.
[0143] In some embodiments, device 101 automatically adjusts the orientation of an object to correspond to another object or surface when the object approaches another object or surface (e.g., regardless of the orientation rules described above for two-dimensional and three-dimensional objects). For example, in Figures 7D-7E, hand 703d is providing indirect movement input to move object 714a to within a threshold distance (e.g., 0.1, 0.5, 1, 3, 6, 12, 24, 36, or 48 cm) of the surface of sofa representation 724a, which is optionally a representation of a physical sofa in the physical environment of device 101. In response, in Figure 7E, device 101 has adjusted the orientation of object 714a to correspond to and / or be parallel to the surface of the approached representation 724a (e.g., optionally such that the bottom surface of object 714a no longer remains parallel to the floor in the physical environment and / or three-dimensional environment 702). 7D-7E, hand 703c is providing indirect movement input to move object 712a to within a threshold distance (e.g., 0.1, 0.5, 1, 3, 6, 12, 24, 36, or 48 cm) of the surface of virtual object 718a. In response, in FIG. 7E, device 101 adjusts the orientation of object 712a to correspond to and / or be parallel to the approached surface of object 718a (e.g., this optionally results in the orientation of object 712a no longer remaining perpendicular to the viewpoint of user 726).
[0144] Additionally, in some embodiments, when an individual object is moved within a threshold distance of a surface of an object (e.g., physical or virtual), device 101 displays a badge on the individual object indicating whether the object is a valid or invalid drop target for the individual object. For example, a valid drop target for an individual object is a drop target to which the individual object can be added and / or a drop target that can contain the individual object, and an invalid drop target for an individual object is a drop target to which the individual object cannot be added and / or a drop target that cannot contain the individual object. In FIG. 7E , object 718a is a valid drop target for object 712a. Accordingly, device 101 displays badge 720 overlaid on the upper right corner of object 712a indicating that object 718a is a valid drop target for object 712a. Further details of valid and invalid drop targets and the associated indications displayed, and other responses of device 101, are described with reference to methods 1000, 1200, 1400, and / or 1600.
[0145] 8A-8K are flowcharts illustrating a method 800 for utilizing different algorithms to move an object in different directions within a three-dimensional environment, according to some embodiments. In some embodiments, method 800 is implemented in a computer system (e.g., computer system 101 of FIG. 1 , such as a tablet, smartphone, wearable computer, or head-mounted device) that includes a display generating component (e.g., display generating component 120 of FIGS. 1, 3, and 4 ) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward in a user's hand (e.g., color sensors, infrared sensors, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 800 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 800 are, optionally, combined, and / or the order of some operations is, optionally, changed.
[0146] In some embodiments, the method 800 is performed in an electronic device (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., 314), such as a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer. In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users. In some embodiments, the one or more input devices include electronic devices or components that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., hand tracking device, hand motion sensor), etc. In some embodiments, the electronic device is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.
[0147] In some embodiments, via a display generating component, a first object, such as object 706a or 708a of FIG. 7A (e.g., an environment corresponding to a physical environment surrounding the display generating component) is displayed in a three-dimensional environment (e.g., a window of an application displayed in the three-dimensional environment, a virtual object (e.g., a virtual clock, a virtual table, etc.) displayed in the three-dimensional environment, etc.). In some embodiments, the three-dimensional environment is generated, displayed, or otherwise made visible by an electronic device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment), and while the first object is selected for movement within the three-dimensional environment, the electronic device, via one or more input devices, responds to movements of individual body parts of a user of the electronic device (e.g., movements of the user's hands and / or arms) within the physical environment in which the display generating component is located, such as input from hands 703a, 703b of FIG. 7A . A corresponding first input is detected 802a (e.g., a pinch gesture of the index finger and thumb of the user's hand while the user's gaze is directed toward the first object while the user's hand is greater than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) from the first object, followed by hand movement in a pinch hand configuration, or a pinch of the index finger and thumb of the user's hand regardless of the location of the user's gaze when the user's hand is less than a threshold distance from the first object, followed by hand movement in a pinch hand configuration). In some embodiments, the first input has one or more characteristics of the input(s) described with reference to methods 1000, 1200, 1400, and / or 1600.
[0148] In some embodiments, in response to detecting 802b a first input, following a determination that the first input includes a movement of a discrete part of the user's body within the physical environment in a first input direction, such as hand 703b in FIG. 7A (e.g., the hand movement in the first input includes (or only includes) movement of the hand away from the user's viewpoint within the three-dimensional environment), the electronic device moves 802c a first object in a first output direction within the three-dimensional environment in accordance with the movement of the discrete part of the user's body within the physical environment in the first input direction, such as movement of object 706a in FIG. 7B (e.g., the user the first object moves away from the user's viewpoint in the three-dimensional environment based on the movement of the user's hand, and the movement of the first object in the first output direction has a first relationship to the movement of a discrete part of the user's body in the physical environment in the first input direction (e.g., the amount and / or manner in which the first object is moved away from the first user's viewpoint is controlled by a first algorithm that translates the movement of the user's hand (e.g., away from the user's viewpoint) into the movement of the first object (e.g., away from the user's viewpoint), exemplary details of which are described below).
[0149] In some embodiments, in accordance with a determination that the first input includes a movement of an individual part of the user's body in the physical environment in a second input direction different from the first input direction, such as hand 703a in FIG. 7A (e.g., the hand movement in the first input includes (or only includes) movement of the hand toward a user's viewpoint in the three-dimensional environment), the electronic device moves the first object in the three-dimensional environment in a second output direction different from the first output direction in accordance with the movement of the individual part of the user's body in the physical environment in the second input direction, such as movement of object 708a in FIG. 7B (e.g., moves the first object toward the user's viewpoint based on the movement of the user's hand), wherein the movement of the first object in the second output direction has a second relationship to the movement of the individual part of the user's body in the physical environment in the second input direction that is different from the first relationship. For example, the amount and / or manner in which a first object is moved toward a first user's viewpoint is controlled by a second algorithm that converts the user's hand movement (e.g., toward the user's viewpoint) into a movement of the first object (e.g., toward the user's viewpoint) that is different from the first algorithm; exemplary details will be described below. In some embodiments, the conversion of hand movement into object movement in the first and second algorithms is different (e.g., different amounts of object movement for a given amount of hand movement). Thus, in some embodiments, movement of the first object in different directions (e.g., horizontal to the user's viewpoint, vertical to the user's viewpoint, further away from the user's viewpoint, toward the user's viewpoint, etc.) is controlled by different algorithms that differently convert the user's hand movement into a movement of the first object in the three-dimensional environment. Using different algorithms to control the movement of an object in a three-dimensional environment for different movement directions allows the device to utilize an algorithm that is more suitable for the movement direction in question to facilitate improved location control for the object in the three-dimensional environment, thereby reducing errors in use and improving user-device interaction.
[0150] In some embodiments, the magnitude of movement of the first object (e.g., object 706a) in the first output direction (e.g., the distance the first object moves in the three-dimensional environment in the first output direction, such as away from the viewpoint as in FIG. 7B ) is independent of the speed of movement of an individual part of the user's body (e.g., hand 703b) in the first input direction (804a) (e.g., the amount of movement of the first object in the three-dimensional environment is based on factors such as the amount of movement of the user's hand providing input in the first input direction, the distance between the user's hand and the user's shoulder when providing input in the first input direction, and / or one or more of various factors described below, but optionally not based on the speed at which the user's hand moves while inputting in the first input direction).
[0151] In some embodiments, the magnitude of movement of the first object (e.g., object 708a) in the second output direction (e.g., the distance the first object moves in the three dimensional environment in the second output direction, such as toward the viewpoint as in FIG. 7B ) is independent of the speed of movement of an individual part of the user's body (e.g., hand 703a) in the second input direction (804b). For example, the amount of movement of the first object in the three dimensional environment is based on factors such as the amount of movement of the user's hand providing input in the second input direction, the distance between the user's hand and the user's shoulder when providing input in the second input direction, and / or one or more of various factors described below, but optionally not based on the speed at which the user's hand moves while inputting in the second input direction. In some embodiments, the distance the object moves in the three dimensional environment is independent of the speed of movement of an individual part of the user, but the speed at which the first object moves through the three dimensional environment is based on the speed of movement of an individual part of the user. Independently activating the amount of object movement from the velocity of individual parts of the user avoids situations in which a sequential input for moving an object that would have returned the object to its original location before the input was received (if the amount of object movement were independent from the velocity of movement of individual parts of the object) results in the object not being returned to its original location (e.g., because the sequential input was provided at a different velocity of individual parts of the user), thereby improving user-device interaction.
[0152] In some embodiments, the (e.g., first input direction and) first output direction is horizontal with respect to a user's viewpoint within the three-dimensional environment (806a), such as with respect to object 712a in FIGS. 7C and 7D (e.g., the electronic device is displaying a viewpoint of the three-dimensional environment associated with a user of the electronic device, and the first input direction and / or the first output direction corresponds to input and / or output for moving the first object horizontally (e.g., substantially parallel to a floor of the three-dimensional environment). For example, the first input direction corresponds to the user's hand moving substantially parallel to a floor (and / or substantially perpendicular to gravity) within the physical environment of the electronic device and / or display generating components, and the first output direction corresponds to movement of the first object substantially parallel to a floor within the three-dimensional environment displayed by the electronic device.
[0153] In some embodiments, the (e.g., second input direction) and second output direction are perpendicular to a user's viewpoint within the three-dimensional environment (806b), such as with respect to object 714a in FIGS. 7C and 7D . For example, the second input direction and / or second output direction correspond to input and / or output for moving a first object vertically (e.g., substantially perpendicular to the floor of the three-dimensional environment). For example, the second input direction corresponds to a user's hand moving substantially perpendicular to the floor (and / or substantially parallel to gravity) within the physical environment of the electronic device and / or display generating components, and the second output direction corresponds to movement of the first object substantially perpendicular to the floor of the three-dimensional environment displayed by the electronic device. Thus, in some embodiments, the amount of movement of the first object for a given amount of movement of a separate part of the user is different for vertical movement of the first object and horizontal movement of the first object. Providing different movement amounts for vertical and horizontal movement inputs allows the device to utilize algorithms that are more suited to the direction of movement in question (e.g., because vertical hand movements may be more difficult for a user to complete than horizontal hand movements due to gravity) to facilitate improved location control over objects within a three-dimensional environment, thereby reducing errors in use and improving user-device interaction.
[0154] In some embodiments, the movement of the individual parts of the user's body within the physical environment in the first input direction and the second input direction has a first magnitude (808a), such as the magnitude of the movement of hands 703a and 703b in FIG. 7A being the same (e.g., the user's hands moving 12 cm within the physical environment of the electronic device and / or display generating components), and the movement of the first object in the first output direction has a second magnitude (808b) that is greater than the first magnitude (e.g., if the input is to move the first object horizontally, the first object moves horizontally within the three-dimensional environment by 12 cm times a first multiplier, such as 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 4, 6, or 10). The movement of the first object in the second output direction has a third magnitude (808c) that is greater than the first magnitude and different from the second magnitude, as shown for object 708a in FIG. 7B (e.g., if the input is to move the first object vertically, the first object moves vertically in the three-dimensional environment by 12 cm times a second multiplier that is greater than the first multiplier, such as 1.2, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 4, 6, or 10). Thus, in some embodiments, an input to move the first object vertically (e.g., up or down) results in more movement of the first object in the three-dimensional environment for a given amount of hand movement compared to an input to move the first object horizontally (e.g., left or right). It should be understood that the above-mentioned multipliers are optionally applied to the entire hand movement or only to the component of the hand movement in each direction (e.g., horizontal or vertical). Providing different movement amounts for vertical and horizontal movement inputs allows the device to utilize algorithms that are more suited to the direction of movement in question (e.g., because vertical hand movements may be more difficult for a user to complete than horizontal hand movements due to gravity) to facilitate improved location control over objects within a three-dimensional environment, thereby reducing errors in use and improving user-device interaction.
[0155] In some embodiments, the first relationship is based on an offset between a second individual part of the user's body (e.g., a shoulder, such as 705a) and an individual part of the user's body (e.g., a hand corresponding to the shoulder, such as 705b), and the second relationship is based on an offset between the second individual part of the user's body (e.g., a shoulder) and an individual part of the user's body (e.g., a hand corresponding to the shoulder) 810. In some embodiments, the offset and / or separation and / or distance and / or angular offset between the user's hand and the user's corresponding shoulder are factors in determining movement of the first object away from and / or towards the user's viewpoint. For example, for a movement of a first object away from the user's viewpoint, the offset between the user's shoulder and hand is optionally recorded at the start of the first input (e.g., at the moment when a pinch down of the index finger and thumb of the hand is detected, such as when the tip of the thumb and the tip of the index finger are detected as touching together before a pinch hand shaped hand movement is detected), and this offset corresponds to a coefficient that is multiplied with the hand movement to determine how far the first object should be moved away from the user's viewpoint. In some embodiments, this coefficient has a value of 1 from 0 to 40 cm (or 5, 10, 15, 20, 30, 50, or 60 cm) of offset between the shoulder and the hand, then increases linearly from 40 (or 5, 10, 15, 20, 30, 50, or 60 cm) to 60 cm (or 25, 30, 35, 40, 50, 70, or 80 cm) of offset, and then increases linearly at a greater rate from 60 cm (or 25, 30, 35, 40, 50, 70, or 80 cm) onwards. For a first object movement towards the user's viewpoint, the offset between the user's shoulder and hand is optionally recorded at the start of the first input (e.g., at the moment a pinch down of the index finger and thumb of the hand is detected, before a pinch-hand shaped hand movement is detected), and the offset is halved. The movement of the hand from the initial offset to the halved offset is optionally set to correspond to the movement of the first object from its current position to the user's viewpoint.Utilizing hand-to-shoulder offset in determining object movement provides a comfortable and consistent object movement response given different starting offsets, thereby reducing errors in use and improving user-device interaction.
[0156] In some embodiments, the (e.g., first input direction and) first output direction correspond to movement 812a away from a user's viewpoint within the three-dimensional environment, as shown by object 706a in FIG. 7B . For example, the electronic device displays a viewpoint of the three-dimensional environment associated with a user of the electronic device, and the first input direction and / or first output direction correspond to input and / or output for moving the first object further away from the user's viewpoint (e.g., substantially parallel to the floor of the three-dimensional environment and / or substantially parallel to the orientation of the user's viewpoint within the three-dimensional environment). For example, the first input direction corresponds to a user's hand moving away from the user's body in the physical environment of the electronic device and / or display generating components, and the first output direction corresponds to movement of the first object away from the user's viewpoint within the three-dimensional environment displayed by the electronic device.
[0157] In some embodiments, the (e.g., second input direction and) second output direction correspond to movement 812b toward a user's viewpoint in the three-dimensional environment, as shown by object 708a in FIG. 7B . For example, the second input direction and / or the second output direction correspond to input and / or output for moving the first object closer to the user's viewpoint (e.g., substantially parallel to the floor of the three-dimensional environment and / or substantially parallel to the orientation of the user's viewpoint within the three-dimensional environment). For example, the second input direction corresponds to a user's hand moving toward the user's body in the physical environment of the electronic device and / or display generating components, and the second output direction corresponds to movement of the first object toward the user's viewpoint within the three-dimensional environment displayed by the electronic device.
[0158] In some embodiments, the movement of the first object in the first output direction is the movement of a discrete part of the user (e.g., movement of hand 703b in FIG. 7A ) in the first input direction increased (e.g., multiplied) by a first value based on the distance between a part of the user (e.g., the user's hand, the user's elbow, the user's shoulder) and a location corresponding to the first object (812c). For example, the amount by which the first object moves in the three-dimensional environment in the first output direction is defined by the amount of movement of the user's hand in the first input direction multiplied by the first value. In some embodiments, the first value is based on the distance between a particular part of the user (e.g., the user's hand, shoulder, and / or elbow) and the location of the first object. For example, in some embodiments, the first value increases as the distance between the first object and the shoulder corresponding to the user's hand providing the input to move the object increases and decreases as the distance between the first object and the user's shoulder decreases. Further details of the first value are provided below.
[0159] In some embodiments, the movement of the first object in the second output direction is the movement of a discrete part of the user (e.g., movement of hand 703a in FIG. 7A ) in the second input direction increased (e.g., multiplied) by a second value different from the first value (812d) that is based on the distance between the user's viewpoint and a location corresponding to the first object in the three-dimensional environment (e.g., and not based on the distance between a part of the user (e.g., the user's hand, the user's elbow, the user's shoulder) and a location corresponding to the first object). For example, the amount by which the first object moves in the three-dimensional environment in the second output direction is defined by the amount of movement of the user's hand in the second input direction multiplied by the second value. In some embodiments, the second value is based on the distance between the user's viewpoint and the location of the first object when the movement of the discrete part of the user in the second input direction is initiated (e.g., upon detecting that the user's hand performs a thumb and index finger pinch gesture while the user's gaze is directed at the first object). For example, in some embodiments, the second value increases as the distance between the first object and the user's viewpoint increases and decreases as the distance between the first object and the user decreases. Further details of the second value are provided below. Providing different multipliers for movement away from and toward the user's viewpoint allows the device to utilize algorithms that are more suited to the direction of movement at issue to facilitate improved location control for objects within the three-dimensional environment (e.g., because the maximum movement toward the user's viewpoint is known (e.g., limited by the movement toward the user's viewpoint), but the maximum movement away from the user's viewpoint may not be known), thereby reducing errors in use and improving user-device interaction.
[0160] In some embodiments, the first value changes as the movement of the individual part of the user in the first input direction (e.g., movement of hand 703b in FIG. 7A ) progresses, and / or the second value changes as the movement of the individual part of the user in the second input direction (e.g., movement of hand 703a in FIG. 7A ) progresses (814). In some embodiments, the first value is a function of the distance of the first object from the user's shoulder (e.g., the first value increases as that distance increases) and the distance of the user's hand from the user's shoulder (e.g., the first value increases as that distance increases). Thus, in some embodiments, as the user's hand moves in the first input direction (e.g., away from the user), the distance of the first object from the shoulder increases (e.g., in response to the movement of the user's hand) and the distance of the user's hand from the user's shoulder increases (e.g., as a result of the movement of the user's hand away from the user's body). Thus, the first value increases. In some embodiments, the second value is additionally or alternatively a function of the distance between the first object and the distance between the user's hand and the user's shoulder at the time of the initial pinch performed by the user's hand, resulting in movement of the user's hand in the second input direction. Thus, in some embodiments, as the user's hand moves in the second input direction (e.g., toward the user's body), the distance of the user's hand from the user's shoulder decreases (e.g., as a result of the user's hand moving toward the user's body). Accordingly, the second value decreases. Providing dynamic multipliers for movement away from and toward the user's viewpoint provides precise location control in a specific range of the object's movement while providing the ability to move the object long distances in other ranges of the object's movement, thereby reducing errors in use and improving user-device interaction.
[0161] In some embodiments, the first value changes in a first manner as the movement of the individual part of the user in the first input direction (e.g., movement of hand 703b in FIG. 7A ) progresses, and the second value changes in a second manner different from the first manner as the movement of the individual part of the user in the second input direction (e.g., movement of hand 703a in FIG. 7A ) progresses (816). For example, the first value changes differently (e.g., by a larger or smaller magnitude and / or by an opposite direction (e.g., an increase or decrease)) as a function of the distance between the first object and the user's hand and / or between the user's hand and the user's shoulder than the second value changes as a function of the distance between the user's hand and the user's shoulder. Providing multipliers that vary differently for movement away from and towards the user's viewpoint accounts for the difference between the user input (e.g., hand movement(s)) required to move an object a potentially unknown distance away from the user's viewpoint and the user input (e.g., hand movement(s)) required to move the object a maximum distance towards the user's viewpoint (e.g., limited by movement to the user's viewpoint), thereby reducing errors in use and improving user-device interaction.
[0162] In some embodiments, the first value remains constant during a given portion of the movement of the user's individual part (e.g., movement of hand 703b in FIG. 7A ) in the first input direction, and the second value does not remain constant during (e.g., any) given portion of the movement of the user's individual part (e.g., movement of hand 703a in FIG. 7A ) in the second input direction (818). For example, the first value is optionally constant within a first distance range between the first object and the user's hand and / or between the user's hand and the user's shoulder (e.g., a relatively short distance, such as a distance less than 5, 10, 20, 30, 40, 50, 60, 80, 100, or 120 cm), and optionally increases linearly as a function of distance for distances longer than the first distance range. In some embodiments, after a threshold distance greater than the first distance range (e.g., after 30, 40, 50, 60, 80, 100, 120, 150, 200, 300, 400, 500, 750, or 1000 cm), the first value is locked to a constant value greater than its value below the threshold distance. In contrast, the second value optionally varies continuously and / or exponentially and / or logarithmically as the distance between the user's hand and the user's shoulder changes (e.g., decreases as the distance between the user's hand and the user's shoulder decreases). Providing multipliers that vary differently for movement away from and towards the user's viewpoint accounts for the difference between the user input (e.g., hand movement(s)) required to move an object a potentially unknown distance away from the user's viewpoint and the user input (e.g., hand movement(s)) required to move the object a maximum distance towards the user's viewpoint (e.g., limited by movement to the user's viewpoint), thereby reducing errors in use and improving user-device interaction.
[0163] In some embodiments, the first multiplier and the second multiplier are based on a ratio of the distance (e.g., between the user's shoulder and a discrete part of the user) to the length of the user's arm (820). For example, the first value is, optionally, a result of multiplying two coefficients. The first coefficient is, optionally, a distance between the first object and the user's shoulder (e.g., optionally as a percentage or proportion of the total arm length corresponding to the hand providing the movement input), and the second coefficient is, optionally, a distance between the user's hand and the user's shoulder (e.g., optionally as a percentage or proportion of the total arm length corresponding to the hand providing the movement input). The second value is, optionally, a result of multiplying two coefficients. The first coefficient is optionally a distance between the first object and the user's viewpoint at the time of an initial pinch gesture performed by the user's hand leading to the movement input provided by the user's hand (e.g., this first coefficient is optionally constant), optionally provided as a percentage or ratio of the total arm length corresponding to the hand providing the movement input, and the second coefficient is optionally a distance between the user's shoulder and the user's hand (e.g., optionally as a percentage or ratio of the total arm length corresponding to the hand providing the movement input). The second factor optionally includes determining the distance between the user's shoulder and hand at the initial pinch gesture performed by the user's hand, and defining that movement of the user's hand to half that initial distance results in the first object moving from its initial / current location all the way to the location of the user's viewpoint. Defining the movement multiplier as being based on a percentage or ratio of the user's arm length (e.g., rather than an absolute distance) enables predictable and consistent device response to input provided by users with different arm lengths, thereby reducing errors in use and improving user-device interaction.
[0164] In some embodiments, as described herein, the electronic device utilizes different algorithms for controlling movement of a first object away from or toward a user's viewpoint (e.g., corresponding to movement of a user's hand away from or toward the user's body, respectively). In some embodiments, the electronic device continuously or periodically detects hand movement (e.g., while remaining in a pinched hand shape), averages a certain number of frames (e.g., 2 frames, 3 frames, 5 frames, 10 frames, 20 frames, 40 frames, or 70 frames) of hand movement detection, and determines whether the user's hand movement corresponds to movement of the first object away from the user's viewpoint during those averaged frames or corresponds to movement of the first object toward the user's viewpoint during those averaged frames. When the electronic device makes these determinations, the electronic device switches between utilizing a first algorithm (e.g., an algorithm for moving the first object away from the user's viewpoint) or a second algorithm (e.g., an algorithm for moving the first object toward the user's viewpoint) that maps the user's hand movement to movement of a first object in a three-dimensional environment. The electronic device optionally dynamically switches between the two algorithms based on the most recent average results of detecting the user's hand movements.
[0165] For the first algorithm, two coefficients are optionally determined at the start of a movement input (e.g., pinching down the user's thumb and index finger) and / or at the start of a movement of the user's hand away from the user's body, and the two coefficients are optionally updated as their components change and / or multiplied with the magnitude of the hand movement to define the magnitude of the resulting object movement. The first coefficient is optionally an object-to-shoulder coefficient corresponding to the distance between the object being moved and the user's shoulder. The first coefficient optionally has a value of 1 for distances between the shoulder and the object from 0 cm to a first distance threshold (e.g., 5 cm, 10 cm, 20 cm, 30 cm, 40 cm, 50 cm, 60 cm, or 100 cm), and optionally has a value that increases linearly as a function of distance up to a maximum coefficient value (e.g., 2, 3, 4, 5, 6, 7.5, 8, 9, 10, 15, or 20). The second coefficient is optionally a shoulder-hand coefficient that has a value of 1 for distances between the shoulder and the hand from 0 cm to a first distance threshold (e.g., 5 cm, 10 cm, 20 cm, 30 cm, 40 cm, 50 cm, 60 cm, or 100 cm, optionally the same as or different from the first distance threshold of the first coefficient), that increases linearly at a first rate as a function of distance from the first distance threshold to a second distance threshold (e.g., 7.5 cm, 15 cm, 30 cm, 45 cm, 60 cm, 75 cm, 90 cm, or 150 cm), and then that increases linearly as a function of distance from the second distance threshold forward at a second rate greater than the first rate. In some embodiments, the first and second coefficients are multiplied together and multiplied with the magnitude of hand movement to determine movement of an object away from the user's viewpoint within the three-dimensional environment. In some embodiments, the electronic device imposes a maximum magnitude on the movement of the object away from the user for a given movement of the user's hand away from the user (e.g., a movement of 1, 3, 5, 10, 30, 50, 100, or 200 meters), and therefore applies a ceiling function having a maximum magnitude to the result of the above multiplication.
[0166] With respect to the second algorithm, a third factor is optionally determined at the start of the movement input (e.g., when the user's thumb and index finger pinch down) and / or at the start of the user's hand movement towards the user's body. The third coefficient is optionally updated as its components change and / or multiplied with the magnitude of the hand movement to define the magnitude of the resulting object movement. The third factor is optionally a shoulder-to-hand factor. An initial distance between the user's shoulder and hand is optionally determined ("initial distance") and recorded at the start of the movement input (e.g., when the user's thumb and index finger pinch down) and / or at the start of the user's hand movement towards the user's body. The value of the third coefficient is defined by a function that maps a movement of the hand that is half the "initial distance" to a movement of the object from its current position in the three-dimensional environment to the user's viewpoint (or to a position corresponding to the initial position of the user's hand when the user's index finger and thumb are pinched down, or to a position offset from that position by a predetermined amount, such as 0.1 cm, 0.5 cm, 1 cm, 3 cm, 5 cm, 10 cm, 20 cm, or 50 cm). In some embodiments, the function has a relatively high value for relatively high shoulder-to-hand distances and a relatively low value (e.g., 1 or greater) for relatively low shoulder-to-hand distances. In some embodiments, the function is a curve that is concave toward lower coefficient values.
[0167] In some embodiments, the above-mentioned distances and / or distance thresholds in the first and / or second algorithms are optionally instead expressed as relative values (e.g., percentages of the user's arm length) rather than absolute distances.
[0168] In some embodiments, the (e.g., first value and / or) second value is based 822 on the position of a respective part of the user (e.g., hand 703a or 703b in FIG. 7A ) when the first object is selected for movement (e.g., detecting an initial pinch gesture performed by the user's hand while the user's gaze is directed at the first object leads to movement of the user's hand to move the first object). For example, as described above, one or more coefficients defining the second value are based on measurements performed by the electronic device corresponding to the distance between the first object and the user's viewpoint at the time of the initial pinch gesture performed by the user's hand and / or the distance between the user's shoulder and the user's hand at the time of the initial pinch gesture performed by the user's hand. Defining one or more of the movement multipliers based on the arm position at the time of selection of the first object for movement enables the device to facilitate the same amount of movement of the first object for a variety of (e.g., multiple, arbitrary) arm positions at which movement input is initiated, rather than making certain movements of the first object unachievable depending on the user's initial arm position at the time of selection of the first object for movement, thereby reducing errors in use and improving user-device interaction.
[0169] In some embodiments, while a first object is selected for movement within the three-dimensional environment, the electronic device detects (824a) via one or more input devices, individual movements of individual parts of the user in a direction horizontal to the user's viewpoint within the three-dimensional environment, such as movement 707c of hand 703c in FIG. 7C . For example, movement of the user's hand left or right relative to gravity. In some embodiments, the hand movements correspond to noise in the user's hand movements (e.g., the user's hand is shaking or trembling). In some embodiments, the hand movements have a velocity and / or acceleration less than a corresponding velocity and / or acceleration threshold. In some embodiments, in response to detecting lateral hand movements having a velocity and / or acceleration greater than the above-mentioned velocity and / or acceleration threshold, the electronic device does not apply noise reduction, as described below, to such movements, but rather moves the first object in accordance with the lateral movement of the user's hand without applying noise reduction.
[0170] In some embodiments, in response to detecting the individual movements of the user's individual parts, the electronic device updates (824b) the location of the first object within the three-dimensional environment based on the noise-reduced individual movements of the user's individual parts, as described with reference to object 712a. In some embodiments, the electronic device moves the first object within the three-dimensional environment according to the noise-reduced magnitude, frequency, velocity, and / or acceleration of the user's hand movements (e.g., using a 1 Eurofilter) rather than according to the non-noise-reduced magnitude, frequency, velocity, and / or acceleration of the user's hand movements. In some embodiments, the electronic device moves the first object with a smaller magnitude, frequency, velocity, and / or acceleration than it would move the first object if noise reduction were not applied. In some embodiments, the electronic device applies such noise reduction to the lateral (and / or lateral component) of the hand movements, but does not apply such noise reduction to the vertical (and / or vertical component) of the hand movements and / or hand movements (and / or components of the hand movements) toward or away from the user's viewpoint. Reducing noise in hand movements relative to lateral hand movements reduces noise in object movements that may be more readily present in side-to-side and / or lateral movements of the user's hand, thereby reducing errors in use and improving user-device interaction.
[0171] In some embodiments, when individual movement of a user's individual part, such as object 712a in FIG. 7C , is detected, in accordance with a determination that a location corresponding to a first object is at a first distance from the user's individual part (e.g., the first object is at a first distance from the user's hand during each lateral movement of the hand), the individual movement of the user's individual part is adjusted based on a first amount of noise reduction used to generate an adjusted movement used to update the location of the first object within the three-dimensional environment, as described with reference to object 712a in FIG. 7C ; and when individual movement of a user's individual part, such as object 716a in FIG. 7C , in accordance with a determination that a location corresponding to the first object is at a second distance less than the first distance from the user's individual part (e.g., the first object is at a second distance from the user's hand during each lateral movement of the hand), the individual movement of the user's individual part is adjusted based on a second amount less than the first amount of noise reduction used to generate an adjusted movement used to update the location of the first object within the three-dimensional environment, as described with reference to object 716a in FIG. 7C . Thus, in some embodiments, the electronic device applies more noise reduction to left-right and / or lateral movements of the user's hand when the movement is directed toward an object farther away from the user's hand than when the movement is directed toward an object closer to the user's hand. Applying different amounts of noise reduction depending on the distance of the object from the user's hand and / or viewpoint allows for a less filtered / more direct response while the object is at a distance where changes to the movement input can be more easily perceived, and a more filtered / less direct response while the object is at a distance where changes to the movement input can less easily be perceived, thereby reducing errors in use and improving user-device interaction.
[0172] In some embodiments, while the first object is selected for movement and during the first input, the electronic device controls the orientation of the first object in the three-dimensional environment in multiple directions (e.g., one or more of pitch, yaw, or roll) according to the corresponding multiple orientation control portions of the first input, such as with respect to object 716a in FIGS. 7C-7D (828). For example, while the user's hand is providing movement input to the first object, input from the user's hand to change the orientation(s) of the first object causes the first object to change its orientation(s) accordingly. For example, input from the hand while the hand is directly manipulating the first object to rotate the first object, tilt the object, etc. (e.g., the hand is closer than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) to the first object during the first input) causes the electronic device to rotate, tilt, etc. the first object according to such input. In some embodiments, such input includes hand rotation, hand tilt, etc. while the hand is providing movement input to the first object. Thus, in some embodiments, during movement input, the first object is free to tilt, rotate, etc. in accordance with the movement input provided by the hand. Changing the orientation of the first object in accordance with the orientation change input provided by the user's hand during movement input allows the first object to be more fully responsive to the input provided by the user, thereby reducing errors in use and improving user-device interaction.
[0173] In some embodiments, while controlling the orientation of the first object in multiple directions (e.g., while the electronic device allows input from a user's hand to control the pitch, yaw, and / or roll of the first object), the electronic device detects (830a) that the first object is within a threshold distance (e.g., 0.1, 0.5, 1, 3, 5, 10, 20, 40, or 50 cm) of a surface in the three-dimensional environment, such as object 712a relative to object 718a or object 714a relative to representation 724a in FIG. 7E . For example, movement input provided by the user's hand moves the first object within the threshold distance of a virtual or physical surface in the three-dimensional environment. For example, the virtual surface is, optionally, a surface of a virtual object in the three-dimensional environment (e.g., a top of a virtual table that is not present in the display generating component and / or the electronic device's physical environment). The physical surface is optionally a surface of a physical object in the physical environment of the electronic device, a representation of which is displayed in the three-dimensional environment by the electronic device (e.g., via a digital pass-through or a physical pass-through, such as through a transparent portion of a display generating component), such as a physical tabletop in the physical environment.
[0174] In some embodiments, in response to detecting that the first object is within a threshold distance of a surface in the three-dimensional environment, the electronic device updates one or more orientations of the first object in the three-dimensional environment (830b) based on the orientation of the surface as described with reference to objects 712a and 714a in FIG. 7E (e.g., not based on the plurality of orientation control portions of the first input). In some embodiments, the orientation of the first object changes to an orientation defined by the surface. For example, if the surface is a wall (e.g., physical or virtual), the orientation of the first object is optionally updated to be parallel to the wall even if no manual input is provided to change the orientation of the first object (e.g., to be parallel to the wall). If the surface is a tabletop (e.g., physical or virtual), the orientation of the first object is optionally updated to be parallel to the tabletop even if no manual input is provided to change the orientation of the first object (e.g., to be parallel to the tabletop). Updating the orientation of the first object based on the orientation of a nearby surface provides a quick way to position the first object relative to the surface in a complementary manner, thereby reducing errors in use and improving user-device interaction.
[0175] In some embodiments, while controlling the orientation of the first object in multiple directions (e.g., while the electronic device allows input from a user's hand to control the pitch, yaw, and / or roll of the first object), the electronic device detects (832a) that the first object is no longer selected for movement, such as with respect to object 716a of FIG. 7E (e.g., detects that the user's hand is no longer in a pinch hand pose with the tip of the index finger touching the tip of the thumb and / or detects that the user's hand has performed a gesture of the tip of the index finger moving away from the tip of the thumb). In some embodiments, in response to detecting that the first object is no longer selected for movement, the electronic device updates (832b) one or more orientations of the first object in the three-dimensional environment to be based on a default orientation of the first object in the three-dimensional environment, as described with reference to object 716a of FIG. 7E . For example, in some embodiments, the default orientation of the object is defined by the three-dimensional environment, such that in the absence of user input to change the object's orientation, the object has a default orientation in the three-dimensional environment. For example, a default orientation of a three-dimensional object in a three-dimensional environment is, optionally, that the bottom surface of such object should be parallel to the floor of the three-dimensional environment. During movement input, the user's hand can optionally provide an orientation-changing input that causes the bottom surface of the object to be non-parallel to the floor. However, upon detecting the end of the movement input, the electronic device optionally updates the orientation of the object so that the bottom surface of the object is parallel to the floor, even if no manual input to change the object's orientation is provided. A two-dimensional object in the three-dimensional environment optionally has a different default orientation. In some embodiments, the default orientation of a two-dimensional object is one in which a normal to the surface of the object is parallel to the orientation of a user's viewpoint in the three-dimensional environment. During movement input, the user's hand can optionally provide an orientation-changing input that causes the normal to the surface of the object to be non-parallel to the orientation of the user's viewpoint.However, upon detecting an end of the movement input, the electronic device optionally updates the orientation of the object so that the normal to the surface of the object is parallel to the orientation of the user's viewpoint, even if no manual input is provided to change the object's orientation. Updating the orientation of the first object to a default orientation ensures that the object does not become unusable over time, thereby reducing errors in use and improving user-device interaction.
[0176] In some embodiments, the first input occurs 834a while a discrete part of the user (e.g., hand 703e in FIG. 7C ) is within a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) of a location corresponding to the first object (e.g., a first input in which the electronic device causes input from the user's hand to control the pitch, yaw, and / or roll of the first object is a direct manipulation input from the hand directed toward the first object). In some embodiments, while the first object is selected for movement, and during a second input corresponding to movement of the first object in the three-dimensional environment, where during the second input a distinct portion of the user is farther than a threshold distance from a location corresponding to the first object (834b), pursuant to a determination that the first object is a two-dimensional object (e.g., the first object is an application window / user interface, a representation of a picture, etc.) (e.g., the second input is an indirect manipulation input from a hand pointed at the first object), the electronic device moves the first object within the three-dimensional environment in accordance with the second input while an orientation of the first object relative to a user's viewpoint within the three-dimensional environment remains constant, as described with reference to object 712a in FIGS. 7C-7D (e.g., in the case of a two-dimensional object, the indirect movement input optionally changes the position of the two-dimensional object within the three-dimensional environment in accordance with the input, but the orientation of the two-dimensional object is controlled by the electronic device such that a normal to the surface of the two-dimensional object remains parallel to the orientation of the user's viewpoint).
[0177] In some embodiments, following a determination that the first object is a three-dimensional object, the electronic device moves the first object within the three-dimensional environment in accordance with the second input (834d) while the orientation of the first object relative to surfaces within the three-dimensional environment remains constant, as described with reference to object 714a in FIGS. 7C-7D (e.g., in the case of a three-dimensional object, the indirect movement input optionally changes the position of the three-dimensional object within the three-dimensional environment in accordance with the input, but the orientation of the three-dimensional object is controlled by the electronic device such that the normal to the bottom surface of the three-dimensional object remains perpendicular to the floor within the three-dimensional environment). Controlling the object's orientation during the movement input ensures that the object does not assume an unusable orientation over time, thereby reducing errors in use and improving user-device interaction.
[0178] In some embodiments, while the first object is selected for movement and during the first input (836a), the electronic device moves the first object within the three-dimensional environment in accordance with the first input (836b) while maintaining the orientation of the first object relative to the user's viewpoint within the three-dimensional environment, as described with reference to object 712a in Figures 7C-7D. For example, during an indirect movement operation of a two-dimensional object, the normal of the object's surface is maintained parallel to the orientation of the user's viewpoint as the object is moved to different positions within the three-dimensional environment.
[0179] In some embodiments, after moving the first object while maintaining the orientation of the first object relative to the user's viewpoint, the electronic device detects (836c) that the first object is within a threshold distance (e.g., 0.1, 0.5, 1, 3, 5, 10, 20, 40, or 50 cm) of a second object in the three-dimensional environment, such as object 712a relative to object 718a in Figure 7E. For example, the first object is moved within a threshold distance of a physical or virtual surface in the three-dimensional environment, such as an application window / user interface, a wall, a table, or a floor surface, as described above.
[0180] In some embodiments, in response to detecting that the first object is within a threshold distance of the second object, the electronic device updates (836d) the orientation of the first object in the three-dimensional environment based on the orientation of the second object, regardless of the orientation of the first object relative to the user's viewpoint, as described with reference to object 712a in FIG. 7E. For example, when the first object is moved within the threshold distance of the second object, the orientation of the first object is no longer based on the user's viewpoint but rather is updated to be defined by the second object. For example, the orientation of the first object is updated to be parallel to the surface of the second object, even if no hand input is provided to change the orientation of the first object. Updating the orientation of the first object based on the orientation of nearby objects provides a quick way to position the first object relative to the object in a complementary manner, thereby reducing errors in use and improving user-device interaction.
[0181] In some embodiments, before the first object is selected for movement within the three-dimensional environment, the first object has a first size (e.g., size within the three-dimensional environment) within the three-dimensional environment (838a). In some embodiments, in response to detecting selection of the first object for movement within the three-dimensional environment (e.g., in response to detecting that the user's hand performs a pinch hand gesture while the user's gaze is directed at the first object), the electronic device scales the first object to have a second size within the three-dimensional environment that is different from the first size (838b), the second size being based, with respect to objects 706a and / or 708a, etc., on the distance between a location corresponding to the first object within the three-dimensional environment and the user's viewpoint when selection of the first object for movement is detected. For example, the objects optionally have a defined size that is optimal or ideal for their current distance from the user's viewpoint (e.g., to ensure that the objects remain interactable by the user at their current distance from the user). In some embodiments, upon detecting an initial selection of an object for movement, the electronic device resizes the object to an optimal or ideal size for the object based on the current distance between the object and the user's viewpoint, without receiving hand input to resize the first object. Further details of such rescaling of objects and / or distance-based sizes are described with reference to method 1000. Updating the size of the first object based on the object's distance from the user's viewpoint provides a quick way to ensure that the object is interactable by the user, thereby reducing errors in use and improving user-device interaction.
[0182] In some embodiments, the first object is selected for movement within the three-dimensional environment in response to detecting a second input including a distinct portion of the user (e.g., hand 703a or 703b) performing a first gesture while the user's gaze is directed toward the first object, followed by maintaining a first shape for a threshold time period (e.g., 0.2, 0.4, 0.5, 1, 2, 3, 5, or 10 seconds) (840). For example, if the first object is currently located within unoccupied space within the three-dimensional environment (e.g., not contained in a container or application window), the selection of the first object for movement is optionally performed in response to a pinch hand gesture performed by the user's hand while the user's gaze is directed toward the first object, followed by the user's hand maintaining the pinch hand shape for the threshold time period. After the second input, moving the hand while maintaining the pinch hand shape optionally moves the first object within the three-dimensional environment in accordance with the hand movement. Selecting the first object for movement in response to a gaze + long pinch gesture provides a quick way to select the first object for movement while avoiding unintentional selection of an object for movement, thereby reducing errors in use and improving user-device interaction.
[0183] In some embodiments, the first object is selected for movement within the three-dimensional environment (842) in response to detecting a second input including movement of a discrete portion of the user (e.g., the user's hand providing the movement input, such as hand 703a) toward the user's viewpoint within the three-dimensional environment greater than a movement threshold (e.g., 0.1, 0.3, 0.5, 1, 2, 3, 5, 10, or 20 cm) while the user's gaze is directed toward the first object. For example, if the first object is contained within another object (e.g., an application window), selection of the first object for movement is optionally in response to a pinch hand gesture made by the user's hand while the user's gaze is directed toward the first object, followed by movement of the user's hand toward the user corresponding to movement of the first object toward the user's viewpoint. In some embodiments, movement of the hand toward the user must correspond to movement exceeding the movement threshold; otherwise, the first object is optionally not selected for movement. After the second input, moving the hand while maintaining the pinch hand shape optionally moves the first object within the three-dimensional environment in accordance with the hand movement. Selecting a first object for movement in response to a gaze+pinch+pull gesture provides a quick way to select a first object for movement while avoiding unintentional selection of an object for movement, thereby reducing errors in use and improving user-device interaction.
[0184] In some embodiments, while the first object is selected for movement and during the first input (844a), the electronic device detects (844b) that the first object is within a threshold distance (e.g., 0.1, 0.5, 1, 3, 5, 10, 20, 40, or 50 cm) of a second object in the three-dimensional environment. For example, the first object is moved within the threshold distance of a physical or virtual surface in the three-dimensional environment, such as an application window / user interface, a wall, a table, or a floor surface, as described above. In some embodiments, in response to detecting that the first object is within a threshold distance of the second object (844c), following a determination that the second object is a valid drop target for the first object (e.g., the second object is a messaging user interface of a messaging application and the first object is capable of containing or accepting the first object, such as a representation of a photo, a video, a text content, or the like, that can be sent to another user via the messaging application), the electronic device displays, via the display generation component, a first visual indication indicating that the second object is a valid drop target for the first object (e.g., does not display a second visual indication described below), as described with reference to object 712a in FIG. 7E. For example, the first object is added to the second object by displaying a badge (e.g., a circle with a + sign) overlaid on an upper right corner of the first object indicating that the second object is a valid drop target and to release the first object at its current location.
[0185] In some embodiments, pursuant to a determination that the second object is not a valid drop target for the first object (e.g., the second object cannot contain or accept the first object, such as when the second object is a messaging user interface of a messaging application and the first object is a user interface of another application or a three-dimensional object), the electronic device, via the display generation component, displays (844e) a second visual indication indicating that the second object is not a valid drop target for the first object, such as by displaying this indication in place of indication 720 of FIG. 7E (e.g., not displaying the first visual indication described above). For example, displaying a badge (e.g., a circle with an X symbol) overlaid on the upper right corner of the first object indicating that the second object is not a valid drop target and releasing the first object at its current location does not add the first object to the second object. Indicating whether a second object is a valid drop target for a first object quickly communicates the results of dropping the first object at its current position, thereby reducing errors during use and improving user-device interaction.
[0186] It should be understood that the particular order in which the operations in method 800 are described is merely exemplary and does not indicate that the described order is the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein.
[0187] 9A-9E show examples of electronic devices that dynamically resize (or not) virtual objects in a three-dimensional environment, according to some embodiments.
[0188] 9A shows electronic device 101 displaying a three-dimensional environment 902 from the perspective of user 926 shown in an overhead view (e.g., facing the back wall of the physical environment in which device 101 is located) via a display generating component (e.g., display generating component 120 of FIG. 1 ). As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes a display generating component (e.g., a touchscreen) and multiple image sensors (e.g., image sensor 314 of FIG. 3 ). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101 could use to capture one or more images of a user or a part of a user (e.g., one or more hands of the user) while the user interacts with electronic device 101. In some embodiments, the user interfaces shown and described below may also be realized on a head-mounted display that includes display generating components that display the user interface or three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).
[0189] 9A , device 101 captures one or more images of the physical environment around device 101 (e.g., operating environment 100), including one or more objects in the physical environment around device 101. In some embodiments, device 101 displays a representation of the physical environment within three-dimensional environment 902. For example, three-dimensional environment 902 includes representation 924a of a sofa (corresponding to sofa 924b in the overhead view), which, optionally, is a representation of a physical sofa within the physical environment.
[0190] 9A, the three-dimensional environment 902 also includes virtual objects 906a (corresponding to object 906b in the overhead view), 908a (corresponding to object 908b in the overhead view), 910a (corresponding to object 910b in the overhead view), 912a (corresponding to object 912b in the overhead view), and 914a (corresponding to object 914b in the overhead view). In FIG. 9A, objects 906a, 910a, 912a, and 914a are two-dimensional objects, and object 908a is a three-dimensional object (e.g., a cube). Virtual objects 906a, 908a, 910a, 912a, and 914a are optionally one or more of an application's user interface (e.g., a messaging user interface, a content browsing user interface, etc.), a three-dimensional object (e.g., a virtual clock, a virtual ball, a virtual car, etc.), or any other element displayed by device 101 that is not included in the device's 101 physical environment. In some embodiments, object 906a is a user interface for playing content (e.g., a video player) and is displayed along with control user interface 907a (corresponding to object 907b in the overhead view). Control user interface 907a optionally includes one or more selectable options for controlling the playback of content presented in object 906a. In some embodiments, control user interface 907a is displayed below and / or slightly in front of object 906a (e.g., closer to the user's viewpoint). In some embodiments, object 908a is displayed with a grabber bar 916a (corresponding to object 916b in the overhead view), which is, optionally, an element to which user-provided input is directed to control the location of object 908a within three-dimensional environment 902. In some embodiments, input is directed at object 908a (and not at grabber bar 916a) to control the location of object 908a within the three-dimensional environment.Thus, in some embodiments, as described in more detail with reference to method 1600, the presence of grabber bar 916a indicates that object 908a can be independently positioned within three-dimensional environment 902. In some embodiments, grabber bar 916a is displayed below and / or slightly in front of (e.g., closer to the user's viewpoint) object 908a.
[0191] In some embodiments, device 101 dynamically scales the size of objects in three dimensional environment 902 as the distance of those objects from the user's viewpoint changes. Whether and / or how much device 101 scales the size of objects is, optionally, based on the type of object (e.g., two-dimensional or three-dimensional) being moved within three dimensional environment 902. For example, in FIG. 9A , hand 903 c is providing movement input to object 906 a to further move object 906 a from the viewpoint of user 926 in three dimensional environment 902, hand 903 b is providing movement input (e.g., directed toward grabber bar 916 a) to object 908 a to further move object 908 a from the viewpoint of user 926 in three dimensional environment 902, and hand 903 a is providing movement input to object 914 a to further move object 914 a from the viewpoint of user 926 in three dimensional environment 902. In some embodiments, such movement input includes moving the user's hand toward or away from the body of user 926 while the user's hand is in a pinch hand configuration (e.g., while the tips of the thumb and index finger of the hand are touching). For example, from FIGS. 9A-9B, device 101 optionally detects that hands 903a, 903b, and / or 903c are moving away from the body of user 926 while in the pinch hand configuration. While multiple hands and corresponding inputs are shown in FIGS. 9A-9E, it should be understood that such hands and inputs need not be detected simultaneously by device 101. Rather, in some embodiments, device 101 responds independently to the hands and / or inputs shown and described in response to independently detecting such hands and / or inputs.
[0192] 9A , device 101 moves objects 906a, 908a, and 914a away from the viewpoint of user 926, as shown in FIG. 9B . For example, device 101 moves object 906a further away from the viewpoint of user 926. In some embodiments, to maintain interactability with the object as it is moved farther away from the user's viewpoint (e.g., by avoiding the displayed size of the object becoming unduly small), device 101 increases the size of the object in three-dimensional environment 902 (e.g., similarly, decreases the size of the object in three-dimensional environment 902 as it is moved closer to the user's viewpoint). However, to avoid user confusion and / or disorientation, device 101 optionally increases the size of the object by an amount that ensures that the object is displayed at successively smaller sizes as it is moved further from the user's viewpoint, although the decrease in the displayed size of the object is, optionally, less than if device 101 had not increased the size of the object in three-dimensional environment 902. Furthermore, in some embodiments, device 101 applies such dynamic scaling to two-dimensional objects, but not to three-dimensional objects.
[0193] 9B , as object 906a is moved further from the viewpoint of user 926, device 101 increased the size of object 906a in three-dimensional environment 902 (e.g., as indicated by the increased size of object 906b in the overhead view compared to the size of object 906b in FIG. 9A ), but increased the size of object 906a in a sufficiently small manner to ensure that the area of the field of view of three-dimensional environment 902 consumed by object 906a decreased from FIG. 9A to FIG. 9B . In this way, interactability with object 906a is optionally maintained as object 906a is moved further from the viewpoint of user 926, while avoiding user confusion and / or disorientation that optionally results from not reducing the displayed size of object 906a as object 906a is moved further from the viewpoint of user 926.
[0194] In some embodiments, controls, such as system controls, displayed with object 906a are not scaled by device 101 or are scaled differently from object 906a. For example, in Figure 9B, control user interface 907a moves with object 906a, away from the viewpoint of user 926, in response to movement input directed at object 906a. However, in Figure 9B, device 101 has increased the size of control user interface 907a within three-dimensional environment 902 (e.g., as reflected in the overhead view) sufficiently so that the display size of control user interface 907a (e.g., the portion of the field of view consumed by control user interface 907a) remains constant even as object 906a and control user interface 907a are moved further from the viewpoint of user 926. 9A, the control user interface 907a had approximately the same width as the object 906a, while in FIG. 9B, the device 101 has increased the size of the control user interface 907a sufficiently so that the width of the control user interface 907a is greater than the width of the object 906a. Thus, in some embodiments, the device 101 increases the size of the control user interface 907a more than it increases the size of the object 906a as the object 906a and the control user interface 907a move farther from the viewpoint of the user 926, and the device 101 optionally similarly decreases the size of the control user interface 907a more than it decreases the size of the object 906a as the object 906a and the control user interface 907a move closer to the viewpoint of the user 926. The device 101 optionally similarly scales the grabber bar 916a associated with the object 908a, as shown in FIG. 9B.
[0195] However, in some embodiments, device 101 does not scale three-dimensional objects in three-dimensional environment 902 as they are moved farther or closer from the viewpoint of user 926. Device 101 optionally does so to mimic the appearance and / or behavior of physical objects as they move closer to or farther from the user in the physical environment. For example, as reflected in the overhead view of three-dimensional environment 902, object 908b remains the same size in FIG. 9B as it was in FIG. 9A. As a result, the displayed size of object 908b is reduced more than the displayed size of object 906a from FIG. 9A to FIG. 9B. Thus, for the same amount of movement of objects 906a and 908a away from user 926's viewpoint, the portion of user 926's field of view consumed by object 908a is, optionally, reduced less than the portion of user 926's field of view consumed by object 906a in FIGS. 9A-9B.
[0196] 9B , object 910a is associated with drop zone 930 for adding an object to object 910a. Drop zone 930 is optionally a volume of space (e.g., a cube or prism) within three-dimensional environment 902 adjacent to and / or in front of object 910a. The boundary or volume of drop zone 930 is optionally not displayed within three-dimensional environment 902, and in some embodiments, the boundary or volume of drop zone 930 is displayed within three-dimensional environment 902 (e.g., by an outline, a volume highlight, a volume shading, etc.). When an object is moved within drop zone 930, device 101 optionally scales the object based on drop zone 930 and / or the object associated with drop zone 930. For example, in FIG. 9B , object 914a has been moved into drop zone 930. As a result, device 101 has scaled object 914a down to fit within drop zone 930 and / or object 910a. The amount by which device 101 scales object 914a is, optionally, different (e.g., different in size and / or different in direction) from the scaling of object 914a performed by device 101 as a function of the distance of object 914a from the viewpoint of user 926 (e.g., as described with reference to object 914a). Thus, in some embodiments, the scaling of object 914a is, optionally, based on the size of object 910a and / or drop zone 930, and optionally not based on the distance of object 914a from the viewpoint of user 926. Furthermore, in some embodiments, device 101 displays a badge or indication 932 overlaid on the top right portion of object 910a indicating whether object 914a is a valid drop target for object 914a.For example, if object 910a is a picture frame container and object 914a is a representation of a picture, object 910a is optionally a valid drop target for object 914a, and indication 932 optionally indicates the same amount, while if object 914a is an application icon, object 910a is optionally an invalid drop target for object 914a, and indication 932 optionally indicates the same amount. Additionally, in Figure 9B, device 101 adjusts the orientation of object 914a to be aligned with object 910a (e.g., parallel to object 910a) in response to object 914a being moved into drop zone 930. If object 914a had not been moved into drop zone 930, object 914a would optionally have a different orientation (e.g., corresponding to the orientation of object 914a in Figure 9A), as described with reference to Figure 9C.
[0197] In some embodiments, when object 914a is removed from drop zone 930, device 101 automatically scales object 906a to a size based on the distance of object 914a from the viewpoint of user 926 (e.g., as described with reference to object 914a) and / or restores the orientation of object 914a to the orientation it had when not aligned with object 910a, as shown in FIG. 9C. In some embodiments, device 101 displays object 910a at this same size and / or orientation if object 914a is an invalid drop target for object 910a, even if object 914a is moved into drop zone 930 and / or in proximity to object 914a. For example, in FIG. 9C, object 914a has been removed from drop zone 930 (e.g., in response to movement input from hand 903a in FIG. 9B) and / or object 910a is not a valid drop target for object 914a. As a result, device 101 displays object 914a at a size in three-dimensional environment 902 that is based on the distance of object 914a from the viewpoint of user 926 (e.g., not based on object 910a), which is optionally larger than the size of object 914a in Figure 9B. Furthermore, device 101 additionally or alternatively displays object 914a at an orientation that is optionally based on orientation input from hand 903a, optionally a different orientation from the orientation of object 914a in Figure 9B, and optionally not based on the orientation of object 910a.
[0198] 9B-9C, hand 903c also provides further movement input directed toward object 906a, causing device 101 to move object 906a further from the perspective of user 926. Device 101 consequently further increases the size of object 906a (e.g., as shown in the overhead view of three-dimensional environment 902), but without exceeding scaling limitations associated with the display size of object 906a, as described above. Device 101 also, optionally, further scales control user interface 907a (e.g., as shown in the overhead view of three-dimensional environment 902) to maintain the display size of control user interface 907a, as described above.
[0199] In some embodiments, device 101 scales some objects as a function of their distance from the viewpoint of user 926 when changes in those distances result from input (e.g., from user 926) to move those objects within three-dimensional environment 902, but device 101 does not scale those objects as a function of their distance from the viewpoint of user 926 when changes in those distances result from a movement of user 926's viewpoint within three-dimensional environment 902 (as opposed to, e.g., movement of the objects within three-dimensional environment 902). For example, in FIG. 9D , user 926's viewpoint is moving toward objects 906a, 908a, and 910a within three-dimensional environment 902 (e.g., corresponding to the user's movement within the physical environment toward the back wall of the room). In response, device 101 updates the display of three-dimensional environment 902 to be from the updated viewpoint of user 926, as shown in FIG. 9D .
[0200] As shown in the overhead view of three-dimensional environment 902, device 101 has not scaled objects 906a, 908a, or 910a within three-dimensional environment 902 as a result of the movement of the viewpoint of user 926 in Figure 9D. The displayed sizes of objects 906a, 908a, and 908c have increased due to the decreased distance between the viewpoint of user 926 and objects 906a, 908a, and 910a.
[0201] 9D , device 101 does not scale object 908a, but rather scales grabber bar 916a (e.g., reduces its size as reflected in the overhead view of three-dimensional environment 902) based on the reduced distance between grabber bar 916a and the viewpoint of user 926 (e.g., to maintain the displayed size of grabber bar 916a). Thus, in some embodiments, in response to a movement of user 926's viewpoint, device 101 does not scale non-system objects (e.g., application user interfaces, representations of content such as pictures or movies, etc.), but scales system objects (e.g., grabber bar 916a, control user interface 907a, etc.) as a function of the distance between user 926's viewpoint and those system objects.
[0202] In some embodiments, in response to device 101 detecting movement input directed at an object that currently has a size that is not based on the distance between the object and the viewpoint of user 926, device 101 scales the object to have a size that is based on the distance between the object and the viewpoint of user 926. For example, in Figure 9D, device 101 detects hand 903b providing movement input directed at object 906a. In response, in Figure 9E, device 101 has scaled down object 906a (e.g., as reflected in the overhead view of three-dimensional environment 902) to a size that is based on the current distance between the viewpoint of user 926 and object 906a.
[0203] Similarly, in response to device 101 detecting the removal of an object from a container object (e.g., a drop target) where the size of the object is based on the container object and not on the distance between the object and the viewpoint of user 926, device 101 scales the object to have a size that is based on the distance between the object and the viewpoint of user 926. For example, in FIG. 9D , object 940a is contained within object 910a and has a size that is based on object 910a (e.g., as described above with reference to objects 914a and 910a). In FIG. 9D , device 101 detects movement input from hand 903a directed toward object 940a (e.g., toward the viewpoint of user 926) to remove object 940a from object 910a. In response, as shown in FIG. 9E , device 101 scales up object 940a (e.g., as reflected in the overhead view of three-dimensional environment 902) to a size that is based on the current distance between the viewpoint of user 926 and object 940a.
[0204] 10A-10I are flowcharts illustrating a method 1000 for dynamically resizing (or not resizing) virtual objects in a three-dimensional environment, according to some embodiments. In some embodiments, method 1000 is performed in a computer system (e.g., computer system 101 of FIG. 1 , such as a tablet, smartphone, wearable computer, or head-mounted device) that includes display generating components (e.g., display generating components 120 of FIGS. 1, 3, and 4 ) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing down the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing camera) or a camera pointing forward from the user's head). In some embodiments, method 1000 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations in method 1000 are, optionally, combined, and / or the order of some operations is, optionally, changed.
[0205] In some embodiments, method 1000 is performed in an electronic device (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., 314), such as a mobile device (e.g., a tablet, smartphone, media player, or wearable device) or a computer. In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users. In some embodiments, the one or more input devices include electronic devices or components that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., hand tracking device, hand motion sensor), etc. In some embodiments, the electronic device is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.
[0206] In some embodiments, the electronic device, via the display generation component, displays 1002a a three-dimensional environment including a first object (e.g., a three-dimensional virtual object such as a car model, or a two-dimensional virtual object such as a user interface of an application on the electronic device) such as object 906a of FIG. 9A at a first location within the three-dimensional environment (e.g., the three-dimensional environment is optionally generated, displayed, or otherwise made viewable by the electronic device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment), wherein the first object has a first size within the three-dimensional environment and occupies a first amount of field of view (e.g., a display size of the first object and / or an angular size of the first object from a current location of the user's viewpoint) from a respective viewpoint (e.g., a viewpoint of a user of the electronic device within the three-dimensional environment), such as the size and display size of object 906a of FIG. 9A . For example, the size of a first object in the three-dimensional environment optionally defines a space / volume that the first object occupies in the three-dimensional environment and is not a function of the distance of the first object from a viewpoint of a user of the device into the three-dimensional environment. The size at which the first object is displayed via the display generating components (e.g., the amount of display area and / or field of view occupied by the first object via the display generating components) is optionally based on the size of the first object in the three-dimensional environment and the distance of the first object from a viewpoint of a user of the device into the three-dimensional environment. For example, a given object having a given size in the three-dimensional environment is optionally displayed via the display generating components at a relatively large size (e.g., occupying a relatively large portion of the field of view from a particular viewpoint) if it is relatively close to the user's viewpoint, and is optionally displayed via the display generating components at a relatively small size (e.g., occupying a relatively small portion of the field of view from a particular viewpoint) if it is relatively far from the user's viewpoint.
[0207] In some embodiments, while displaying a three-dimensional environment including a first object at a first location within the three-dimensional environment, the electronic device receives (1002b) via one or more input devices a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment, such as an input from hand 903c directed at object 906a in Figure 9A. For example, a pinch gesture of the index finger and thumb of a user's hand while the user's gaze is directed toward the first object while the user's hand is greater than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) from the first object, followed by movement of the hand in a pinch hand shape, or a pinch of the index finger and thumb of a user's hand regardless of the location of the user's gaze when the user's hand is less than the threshold distance from the first object, followed by movement of the hand in a pinch hand shape. The hand movement is optionally away from the user's viewpoint, which optionally corresponds to a request to move a first object within the three-dimensional environment away from the user's viewpoint. In some embodiments, the first input has one or more characteristics of the input(s) described with reference to methods 800, 1200, 1400, and / or 1600.
[0208] In some embodiments, in response to receiving a first input (1002c), and following a determination (1002d) that the first input corresponds to a request to move the first object away from the individual viewpoint (e.g., a request to move the first object further away from a location in the three dimensional environment where the electronic device is displaying the three dimensional environment), such as an input from hand 903c pointed at object 906a in FIG. 9A , the electronic device moves the first object from the first location away from the individual viewpoint to a second location in the three dimensional environment in accordance with the first input (1002e), the second location being farther from the individual viewpoint than the first location, as shown for object 906a in FIG. 9B (e.g., if the first input corresponds to an input to move the first object from a location 10 meters from the user's viewpoint to a location 20 meters from the user's viewpoint, the first object in the three dimensional environment is moved from a location 10 meters from the user's viewpoint to a location 20 meters from the user's viewpoint). In some embodiments, the electronic device scales (1002f) the first object so that, when the first object is located at the second location, the first object has a second size in the three-dimensional environment that is larger than the first size (e.g., increasing the size of the first object in the three-dimensional environment as the first object moves farther from the user's viewpoint, and optionally decreasing the size of the first object in the three-dimensional environment as the first object moves closer to the user's viewpoint) and occupies a second amount of the field of view from the respective viewpoint, the second amount being smaller than the first amount, as shown for object 906a in FIG. 9B . For example, the display generating component may increase the size of the first object in the three-dimensional environment as the first object moves away from the user's viewpoint by an amount that keeps the display area occupied by the first object the same or by an amount that increases the size of the first object as the first object moves away from the user's viewpoint.In some embodiments, the first object is increased in size within the three-dimensional environment as it moves further from the user's viewpoint so as to maintain the user's ability to interact with the first object; the first object, optionally, would become too small to interact with if not scaled up as it moves further from the user's viewpoint. However, to avoid the sensation that the first object is not actually moving further from the user's viewpoint (e.g., which optionally occurs if the displayed size of the first object remains the same or increases as the first object moves further from the user's viewpoint), the amount of scaling of the first object performed by the electronic device is low enough to ensure that the displayed size of the first object (e.g., the amount of field of view from individual viewpoints occupied by the first object) decreases as the first object moves further from the user's viewpoint. Scaling the first object as a function of the distance of the first object from the user's viewpoint while maintaining the display or angular size of the first object increasing (when the first object is moving towards the user's viewpoint) or decreasing (when the first object is moving away from the user's viewpoint) ensures continued interactability with the first object at a range of distances from the user's viewpoint while avoiding disrupting the presentation of the first object within the three-dimensional environment, thereby improving user-device interaction.
[0209] In some embodiments, while receiving the first input, in accordance with determining that the first input corresponds to a request to move the first object away from the individual viewpoint, the method continuously scales 1004 the first object to an increasing size (e.g., larger than the first size) as the first object moves further from the individual viewpoint, as described with reference to object 906a in FIG. 9B . Thus, in some embodiments, the change in size of the first object occurs continuously as the distance of the first object from the individual viewpoint changes (e.g., the size of the first object is scaled down because the first object is approaching the individual viewpoint, or the size of the first object is enlarged because the first object is further away from the individual viewpoint). In some embodiments, the size of the first object that increases as the first object moves further from the individual viewpoint is the size of the first object in the three-dimensional environment, which is, optionally, an amount different from the size of the first object in the user's field of view from the individual viewpoint, as described in further detail below. Continuously scaling the first object as the distance between the first object and the respective viewpoint changes provides the user with immediate feedback about the movement of the first object, thereby improving user-device interaction.
[0210] In some embodiments, the first object is of a first type, such as object 906a, that is a two-dimensional object (e.g., the first object is a two-dimensional object in the three-dimensional environment, such as a user interface of a messaging application for messaging other users), and the three-dimensional environment further includes 1006a a second object that is of a second type different from the first type, such as object 908a, that is a three-dimensional object (e.g., the second object is a three-dimensional object, such as a virtual three-dimensional representation of a car, a building, a clock, etc. in the three-dimensional environment). In some embodiments, while displaying the three-dimensional environment including the second object at a third location within the three-dimensional environment, the second object has a third size within the three-dimensional environment and occupies a third amount of the field of view from the respective viewpoint, and the electronic device receives 1006b, via one or more input devices, a second input corresponding to a request to move the second object away from the third location within the three-dimensional environment, such as an input from hand 903b pointed at object 908a in FIG. For example, a pinch gesture of the index finger and thumb of the user's hand followed by hand movement in a pinch hand shape while the user's gaze is directed toward the second object when the user's hand is more than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) from the third object, or a pinch gesture of the index finger and thumb of the user's hand followed by hand movement in a pinch hand shape when the user's hand is less than a threshold distance from the third object, regardless of the location of the user's gaze. The hand movement is optionally in a direction away from the user's viewpoint, which optionally corresponds to a request to move a third object within the three-dimensional environment away from the user's viewpoint, or the hand movement is in a direction towards the user's viewpoint, which optionally corresponds to a request to move a third object towards the user's viewpoint. In some embodiments, the second input has one or more of the characteristics of the input(s) described with reference to methods 800, 1200, 1400, and / or 1600.
[0211] In some embodiments, in response to receiving a third input, following a determination that the second input corresponds to a request to move the second object away from the individual viewpoint (e.g., a request to move the third object further away from a location in the three-dimensional environment where the electronic device is displaying the three-dimensional environment), such as an input from hand 903b pointed at object 908a in FIG. 9A (1006c), the electronic device moves the second object from the third location in the three-dimensional environment away from the individual viewpoint (1006d) in accordance with the second input, the fourth location being farther from the individual viewpoint than the third location without scaling the second object such that when the second object is located at the fourth location, the second object has a third size in the three-dimensional environment and occupies a fourth amount less than the third amount of the field of view from the individual viewpoint, as shown for object 908a in FIG. 9B. For example, the three-dimensional object is optionally not scaled based on distance from the user's viewpoint (as compared to a two-dimensional object that is optionally scaled based on distance from the user's viewpoint). Thus, when the three-dimensional object is moved farther from the user's viewpoint, the amount of field of view occupied by the three-dimensional object is optionally reduced, and when the three-dimensional object is moved closer to the user's viewpoint, the amount of field of view occupied by the three-dimensional object is optionally increased. Scaling two-dimensional objects rather than three-dimensional objects treats the three-dimensional objects more similar to physical objects, which results in behavior that is familiar to and expected by users, thereby improving user-device interaction and reducing errors in use.
[0212] In some embodiments, the second object is displayed (1008a) with a control user interface for controlling one or more actions associated with the second object, such as object 907a or 916a in FIG. 9A . For example, the second object is displayed with a selectable and movable user interface element to move the second object within the three-dimensional environment in a manner corresponding to the movement. For example, the user interface element is, optionally, a grabber bar displayed below the second object and graspable to move the second object within the three-dimensional environment. The control user interface optionally is or includes one or more of a grabber bar, a selectable option selectable to discontinue displaying the second object within the three-dimensional environment, a selectable option selectable to share the second object with another user, etc.
[0213] In some embodiments, when the second object is displayed at a third location, the control user interface is displayed at the third location and has a fourth size in the three-dimensional environment (1008b). In some embodiments, when the second object is displayed at a fourth location (e.g., in response to a second input to move the second object), the control user interface is displayed at the fourth location and has a fifth size in the three-dimensional environment that is larger than the fourth size, as shown for objects 907a and 916a (1008c). For example, the control user interface moves along with the second object in accordance with the same second input. In some embodiments, even if the second object is not scaled in the three-dimensional environment based on the distance between the second object and the user's viewpoint, the control user interface displayed with the second object is scaled in the three-dimensional environment based on the distance between the control user interface and the user's viewpoint (e.g., to ensure continued interactability of the control user interface elements by the user). In some embodiments, the control user interface elements are scaled in the same manner as the first object is scaled based on movement toward / away from the user's viewpoint. In some embodiments, the control user interface is scaled smaller or larger than how the first object is scaled based on movement toward / away from the user's viewpoint. In some embodiments, the control user interface is scaled such that when the second object is moved toward / away from the user's viewpoint, the amount of the user's field of view occupied by the control user interface does not change. Scaling the control user interface of a three-dimensional object ensures that the user can interact with the control user interface elements regardless of the distance between the three-dimensional object and the user's viewpoint, thereby improving user-device interaction and reducing errors during use.
[0214] In some embodiments, while displaying a three-dimensional environment including a first object at a first location within the three-dimensional environment, the first object has a first size within the three-dimensional environment, the distinct viewpoint is a first viewpoint, and the electronic device detects (1010a) a movement of a user's viewpoint from the first viewpoint to a second viewpoint, changing a distance between the user's viewpoint and the first object, such as a movement of the viewpoint of user 926 in FIG. 9D . For example, the user moves within the physical environment of the electronic device and / or provides input to the electronic device to move the user's viewpoint from the first distinct location to a second distinct location within the three-dimensional environment, such that the electronic device displays the three-dimensional environment from the user's updated viewpoint. The movement of the viewpoint optionally moves the viewpoint closer to or farther from the first object compared to the distance when the viewpoint was at the first viewpoint.
[0215] In some embodiments, in response to detecting a movement of the viewpoint from the first viewpoint to the second viewpoint, the electronic device updates (1010b) the display of the three-dimensional environment to be from the second viewpoint without scaling the size of the first object at the first location in the three-dimensional environment, as described with reference to objects 906a, 908a, and 910a in Figure 9D. For example, the first object remains the same size in the three-dimensional environment as when the viewpoint was at the first viewpoint, but the amount of field of view occupied by the first object when the viewpoint is at the second viewpoint is, optionally, larger (if the viewpoint has moved closer to the first object) or smaller (if the viewpoint has moved farther from the first object) than the amount of field of view occupied by the first object when the viewpoint was at the first viewpoint. Forgoing scaling a first object in response to changes in the distance between the first object and the user's individual viewpoint as a result of viewpoint movement (as opposed to object movement) ensures that changes in the three-dimensional environment occur when expected (e.g., in response to user input), reducing user disorientation and thereby improving user-device interaction and reducing errors in use.
[0216] In some embodiments, the first object is an object of a first type (e.g., a two-dimensional user interface of an application on the electronic device, a content object such as a three-dimensional representation of an object such as a car, or more generally, an object that is content or corresponds to a system (e.g., operating system) user interface of the electronic device rather than an object that corresponds to or is a system (e.g., operating system) user interface of the electronic device), and the three-dimensional environment further includes a second object (1012a) that is an object of a second type different from the first type, such as object 916a of FIG. 9C (e.g., a control user interface for the individual object as described above, such as a grabber bar for moving the individual object within the three-dimensional environment). In some embodiments, while displaying a three-dimensional environment including the second object at a third location within the three-dimensional environment (e.g., displaying a grabber bar for a distinct object also at the third location within the three-dimensional environment), where the second object has a third size within the three-dimensional environment and the user's viewpoint is the first viewpoint, the electronic device detects (1012b) a movement of the viewpoint from a first viewpoint to a second viewpoint that changes the distance between the user's viewpoint and the second object, such as a movement of the viewpoint of user 926 in FIG. 9D . For example, the user moves within the electronic device's physical environment and / or provides input to the electronic device to move the user's viewpoint from the first distinct location to the second distinct location within the three-dimensional environment, such that the electronic device displays the three-dimensional environment from the user's updated viewpoint. The movement of the viewpoint optionally moves the viewpoint closer to or farther from the second object compared to the distance when the viewpoint was at the first viewpoint.
[0217] In some embodiments, in response to detecting a movement of the respective viewpoints (1012c), the electronic device updates the display of the three-dimensional environment to be from a second viewpoint (1012d), as shown in FIG. 9D. In some embodiments, the electronic device scales the size of the second object at the third location to a fourth size in the three-dimensional environment that is different from the third size, such as scaling object 916a in FIG. 9D (1012e). For example, the second object is scaled in response to a movement of the user's viewpoint based on the updated distance between the user's viewpoint and the second object. In some embodiments, if the user's viewpoint moves closer to the second object, the second object is reduced in size in the three-dimensional environment, and if the user's viewpoint moves away from the user's viewpoint, the second object is enlarged in size. The amount of field of view occupied by the second object optionally remains constant or increases or decreases in the manner described herein. Scaling some types of objects in response to movement of the user's viewpoint ensures that the user is able to interact with the object regardless of the distance between the object and the user's viewpoint, even if the change in distance is due to movement of the viewpoint, thereby improving user-device interaction and reducing errors in use.
[0218] In some embodiments, while displaying a three-dimensional environment including a first object at a first location within the three-dimensional environment, the electronic device detects (1014a) a movement of a user's viewpoint within the three-dimensional environment from a first viewpoint to a second viewpoint, changing the distance between the viewpoint and the first object, such as a movement of the viewpoint of user 926 in FIG. 9D . For example, the user moves within the electronic device's physical environment and / or provides input to the electronic device to move the user's viewpoint to the second viewpoint within the three-dimensional environment, such that the electronic device displays the three-dimensional environment from the user's updated viewpoint. The movement of the viewpoint optionally moves the viewpoint closer to or farther from the first object compared to the distance when the viewpoint was at the first viewpoint.
[0219] In some embodiments, in response to detecting a movement of the viewpoint, the electronic device updates (1014b) the display of the three-dimensional environment to be from the second viewpoint without scaling the size of a first object at a first location in the three-dimensional environment, as shown by objects 906a, 908a, or 910a in Figure 9D. The first object is, optionally, not scaled in response to the movement of the viewpoint, as described above.
[0220] In some embodiments, while displaying the first object at a first location within the three-dimensional environment from a second perspective, the electronic device receives (1014c) a second input via one or more input devices corresponding to a request to move the first object from the first location within the three-dimensional environment to a third location within the three-dimensional environment that is farther from the second distinct location than the first location, such as a movement input directed toward object 906a in FIG. 9D. For example, a pinch gesture of the index finger and thumb of the user's hand while the user's gaze is directed toward the first object while the user's hand is more than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) from the first object, followed by movement of the hand in a pinch-hand shape, or a pinch of the index finger and thumb of the user's hand regardless of the location of the user's gaze when the user's hand is less than a threshold distance from the first object, followed by movement of the hand in a pinch-hand shape. The hand movement is optionally away from the user's viewpoint, which optionally corresponds to a request to move a first object within the three-dimensional environment away from the user's viewpoint. In some embodiments, the second input has one or more characteristics of the input(s) described with reference to methods 800, 1200, 1400, and / or 1600.
[0221] In some embodiments, while detecting the second input, before moving the first object away from the first location (e.g., in response to detecting a pinch down of the user's index finger and thumb, e.g., when the tip of the thumb and the tip of the index finger are detected touching together, before detecting a movement of the hand while maintaining a pinched hand shape), the electronic device scales (1014d) the size of the first object to be a third size different from the first size based on the distance between the first object and a second viewpoint when the start of the second input was detected, such as scaling object 906a in FIG. 9E before object 906a was moved (e.g., when a pinch down of the user's index finger and thumb was detected). If the viewpoint is moved to a location closer to the first object, the first object is optionally scaled down in size, and if the viewpoint is moved to a location farther away from the first object, the first object is optionally enlarged in size. The amount of scaling of the first object (and / or the resulting amount of field of view of the viewpoint occupied by the first object) is, optionally, as described previously with reference to the first object. Thus, in some embodiments, even if the first object is not scaled in response to movement of the user's viewpoint, upon detecting the start of a subsequent movement input, it is scaled based on the current distance between the first object and the user's viewpoint. Scaling the first object upon detecting a movement input ensures that the first object is appropriately sized for its current distance from the user's viewpoint, thereby improving user-device interaction.
[0222] In some embodiments, the three-dimensional environment further includes a second object at a third location within the three-dimensional environment (1016a). In some embodiments, in response to receiving the first input (1016b), in accordance with determining that the first input corresponds to a request to move the first object to a fourth location within the three-dimensional environment (e.g., a location within the three-dimensional environment that does not include another object), the fourth location being the first distance from the respective viewpoint, the electronic device displays the first object at the fourth location within the three-dimensional environment (1016c), the first object having a third size within the three-dimensional environment. For example, the first object is scaled based on the first distance, as previously described.
[0223] In some embodiments, following a determination that the first input satisfies one or more criteria, including individual criteria that are met when the first input corresponds to a request to move the first object to a third location within the three-dimensional environment, such as an input directed at object 914a of FIG. 9A (e.g., moving the first object to a second object, where the second object is a valid drop target for the first object. Valid drop targets and invalid drop targets are described in more detail with reference to methods 1200 and / or 1400), the third location, a first distance from a individual viewpoint (e.g., the second object is at the same distance from the user's viewpoint as the distance of the fourth location from the user's viewpoint), the electronic device displays (1016d) the first object at the third location within the three-dimensional environment, wherein the first object has a fourth size within the three-dimensional environment that is different from the third size, as shown for object 914a of FIG. 9B. In some embodiments, when the first object is moved to an object (e.g., a window) that is a valid drop target for the first object, or within a threshold distance, such as 0.1, 0.2, 0.5, 1, 2, 3, 5, 10, 20, 30, or 50 cm, of the object, the electronic device scales the first object differently than it scales the first object when the first object is not moved to the object (e.g., scales the first object not based on the distance between the user's viewpoint and the first object). Thus, even though the first object is still the first distance from the user's viewpoint when the first object is moved to the third location, the first object has a different size and therefore occupies a different amount of the user's field of view than when the first object is moved to the fourth location.Scaling the first object differently when it is moved to another object provides visual feedback to the user that the first object has been moved to another object that is potentially a drop target / container for the first object, thereby improving user-device interaction and reducing errors in use.
[0224] In some embodiments, the fourth size of the first object is based on the size of the second object (1018), as shown by object 914a in FIG. 9B . For example, the first object is sized to fit within the second object. If the second object is a user interface for a messaging application and the first object is a representation of a photo, the first object is optionally scaled up or down, for example, to be an appropriate size for inclusion / display within the second object. In some embodiments, the first object is scaled to be a particular percentage (e.g., 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, or 70%) of the size of the second object. Scaling the first object based on the size of the second object ensures that the first object is appropriately sized relative to the second object (e.g., not too large so as to obstruct the second object, and not too small so as to be appropriately visible and / or interactable within the second object), thereby improving user-device interaction and reducing errors in use.
[0225] In some embodiments, while the first object is at a third location within the three-dimensional environment and has a fourth size based on the size of the second object (e.g., while the first object is contained within the second object, such as a representation of a photo contained within a photo browsing and viewing user interface), the electronic device receives (1020a) via one or more input devices a second input corresponding to a request to move the first object away from the third location within the three-dimensional environment, such as with respect to object 940a in FIG. 9D (e.g., an input removing the first object from the second object, such as moving the first object beyond a threshold distance, such as 0.1, 0.2, 0.5, 1, 2, 5, 10, 20, or 30 cm, from the second object). For example, a pinch gesture of the index finger and thumb of the user's hand while the user's gaze is directed toward the first object while the user's hand is more than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) from the first object, followed by movement of the hand in a pinch hand shape, or a pinch of the index finger and thumb of the user's hand regardless of the location of the user's gaze when the user's hand is less than a threshold distance from the first object, followed by movement of the hand in a pinch hand shape. The hand movement is optionally toward the user's viewpoint, which optionally corresponds to a request to move the first object toward the user's viewpoint (e.g., and / or away from the second object) within the three-dimensional environment. In some embodiments, the second input has one or more characteristics of the input(s) described with reference to methods 800, 1200, 1400, and / or 1600.
[0226] In some embodiments, in response to receiving the second input, the electronic device displays the first object at a fifth size (1020b), where the fifth size is not based on the size of the second object, as shown by object 940a in FIG. 9E. In some embodiments, the electronic device immediately scales the first object when it is removed from the second object (e.g., before detecting movement of the first object away from the third location). In some embodiments, the size to which the electronic device scales the first object is based on the distance between the first object and the user's viewpoint (e.g., at the moment the first object is removed from the second object) and is no longer based on or proportional to the size of the second object. In some embodiments, the fifth size is the same size as the first object was displayed at immediately before it reached / added to the second object during the first input. Scaling the first object when it is removed from the second object not based on the size of the second object ensures that the first object is appropriately sized for its current distance from the user's viewpoint, thereby improving user-device interaction and reducing errors in use.
[0227] In some embodiments, the individual criteria are satisfied (1022) when the first input corresponds to a request to move the first object anywhere within a volume within the three-dimensional environment that includes the third location, such as within volume 930 of FIG. 9B . In some embodiments, the drop zone for the second object is a volume within the three-dimensional environment (e.g., its boundary, optionally, is not displayed within the three-dimensional environment) that encompasses at least a portion of, but not all, the second object. In some embodiments, the volume extends from a surface of the second object toward a user's viewpoint. In some embodiments, moving the first object anywhere within the volume causes the first object to scale based on the size of the second object (e.g., rather than based on the distance between the first object and the user's viewpoint). In some embodiments, detecting the end of the first input while the first object is within the volume (e.g., detecting the release of a pinch hand shape by the user's hands) causes the first object to be added to the second object, as described with reference to methods 1200 and / or 1400. Providing a volume in which a first object is scaled based on a second object facilitates easier interaction between the first object and the second object, thereby improving user-device interaction and reducing errors in use.
[0228] In some embodiments, in accordance with determining, while receiving the first input (e.g., before detecting the end of the first input, as described above), that the first object has moved to a third location in accordance with the first input and that one or more criteria have been met, the electronic device changes (1024) the appearance of the first object to indicate that the second object is a valid drop target for the first object, as described with reference to object 914a in FIG. 9B . For example, changing the size of the first object, changing the color of the first object, changing the translucency or brightness of the first object, displaying a badge with a “+” symbol overlaid on the upper right corner of the first object, and / or the like, to indicate that the second object is a valid drop target for the first object. The change in the appearance of the first object is, optionally, as described with reference to methods 1200 and / or 1400 in connection with a valid drop target. Changing the appearance of the first object to indicate that it has been moved to a valid drop target provides the user with visual feedback that the first object will be added to the second object when the user finishes the movement input, thereby improving user-device interaction and reducing errors during use.
[0229] In some embodiments, the one or more criteria include a criterion that is met when the second object is a valid drop target for the first object and that is not met when the second object is not a valid drop target for the first object (1026a) (e.g., examples of valid and invalid drop targets are described with reference to methods 1200 and / or 1400). In some embodiments, in response to receiving the first input (1026b), following a determination that the respective criterion is met but the first input does not satisfy the one or more criteria because the second object is not a valid drop target for the first object (e.g., the first object has been moved to a location that would allow the first object to be added to the second object if the second object were a valid drop target for the first object), the electronic device displays the first object at a fourth location within the three-dimensional environment (1026c), wherein the first object has a third size within the three-dimensional environment, as shown by object 914a in FIG. 9C. For example, because the second object is not a valid drop target for the first object, the first object is scaled based on the current distance between the user's viewpoint and the first object rather than being scaled based on the size of the second object. Rescinding the scaling of the first object based on the second object provides the user with visual feedback that the second object is not a valid drop target for the first object, thereby improving user-device interaction and reducing errors in use.
[0230] In some embodiments, in response to receiving the first input (1028a), following a determination that the first input satisfies one or more criteria (e.g., the first object has been moved to a drop location for the second object and the second object is a valid drop target for the first object), the electronic device updates (1028b) the orientation of the first object relative to the respective viewpoint based on the orientation of the second object relative to the respective viewpoint, as shown for object 914a relative to object 910a in Figure 9B. Additionally or alternatively, in addition to or instead of scaling the first object when it is moved to a valid drop target, the electronic device updates / changes the pitch, yaw, and / or roll of the first object to align with the orientation of the second object. For example, if the second object is a planar object or has a planar surface (e.g., is a three-dimensional object with a planar surface), when the first object is moved toward the second object, the electronic device changes the orientation of the first object so that the first object (or, if the first object is a three-dimensional object, a surface thereon) is parallel to the second object (or the surface of the second object). Changing the orientation of the first object when it is moved toward the second object ensures that the first object, if dropped within the second object, is properly positioned within the second object, thereby improving user-device interaction.
[0231] In some embodiments, the three-dimensional environment further includes a second object at a third location within the three-dimensional environment (1030a). In some embodiments, while receiving the first input (1030b), in accordance with a determination that the first input corresponds to a request to move the first object through the third location and farther from the individual viewpoint than the third location (e.g., an input corresponding to moving the first object through the second object, as described with reference to method 1200) (1030c), the electronic device moves the first object from the first location to the third location away from the individual viewpoint (1030d) while scaling the first object within the three-dimensional environment based on the distance between the individual viewpoint and the first object in accordance with the first input, e.g., moving and scaling object 914a from FIG. 9A to FIG. 9B until it reaches object 910a. For example, the first object is freely moved backward away from the user's viewpoint in accordance with the first input until it reaches the second object, as described with reference to method 1200. Because the first object is moving further away from the user's viewpoint as it moves backward toward the second object, the electronic device optionally scales the first object based on the current changing distance between the first object and the user's viewpoint, as described above.
[0232] In some embodiments, after the first object reaches the third location, the electronic device maintains the display of the first object at the third location (1030e) without scaling the first object while continuing to receive the first input, such as when input from the hand 903a directed at the object 914a corresponds to continued movement through the object 910a while the object 914a remains within the volume 930 of FIG. 9B . For example, similar to that described with reference to method 1200, the first object resists movement through the second object upon colliding with and / or reaching the second object, even if further input from the user's hand to move the first object through the second object is detected. Because the distance between the user's viewpoint and the first object has not changed while the first object remains at the second / third location, the electronic device stops scaling the first object within the three-dimensional environment in accordance with further input to move the first object through / over the second object. When an input is received through the second object that is large enough to break through the second object (e.g., as described with reference to method 1200), the electronic device optionally resumes scaling the first object based on the current changing distance between the first object and the user's viewpoint. Ceasing to scale the first object when the first object is pinned relative to the second object provides feedback to the user that the first object is no longer moving within the three-dimensional environment, thereby improving user-device interaction.
[0233] In some embodiments, scaling the first object follows a determination that a second amount of the field of view from the distinct viewpoint occupied by the first object at the second size is greater than a threshold amount of the field of view (1032a) (e.g., the electronic device scales the first object based on a distance between the first object and the user's viewpoint so long as the amount of the user's field of view occupied by the first object is greater than a threshold amount, such as 0.1%, 0.5%, 1%, 3%, 5%, 10%, 20%, 30%, or 50% of the field of view). In some embodiments, while displaying the first object at a distinct size within the three-dimensional environment, where the first object occupies a first distinct amount of the field of view from the distinct viewpoint, the electronic device receives a second input via one or more input devices corresponding to a request to move the first object away from the distinct viewpoint (1032b). For example, a pinch gesture of the index finger and thumb of the user's hand while the user's gaze is directed toward the first object while the user's hand is more than a threshold distance (e.g., 0.2, 0.5, 1, 2, 3, 5, 10, 12, 24, or 26 cm) from the first object, followed by movement of the hand in a pinch hand shape, or a pinch of the index finger and thumb of the user's hand regardless of the location of the user's gaze when the user's hand is less than a threshold distance from the first object, followed by movement of the hand in a pinch hand shape. The hand movement is optionally away from the user's viewpoint, which optionally corresponds to a request to move the first object within the three-dimensional environment away from the user's viewpoint. In some embodiments, the second input has one or more characteristics of the input(s) described with reference to methods 800, 1200, 1400, and / or 1600.
[0234] In some embodiments, in response to receiving the second input (1032c), following a determination that the first individual amount of field of view from the individual viewpoint is less than a threshold amount of field of view (e.g., the amount of field of view occupied by the first object has reached the threshold amount of field of view and / or the first object has been moved to a distance from the user's viewpoint in the three-dimensional environment that is less than the threshold amount), the electronic device moves the first object away from the individual viewpoint in accordance with the second input (1032d) without scaling the size of the first object in the three-dimensional environment, such as when the device 101 stops scaling object 906a from FIG. 9B to FIG. 9C . In some embodiments, the electronic device no longer scales the size of the first object in the three-dimensional environment based on the distance between the first object and the user's viewpoint when the first object is sufficiently far from the user's viewpoint such that the amount of field of view occupied by the first object is less than the threshold amount. In some embodiments, at this point, the size of the first object remains constant in the three-dimensional environment as the first object continues to move further away from the user's viewpoint. In some embodiments, if the first object is subsequently moved closer to the user's viewpoint such that the amount of field of view occupied by the first object reaches and / or exceeds a threshold amount, the electronic device resumes scaling the first object based on the distance between the first object and the user's viewpoint. Ceasing to scale the first object when it consumes less than the user's threshold field of view conserves processing resources of the device when interaction with the first object is not active, thereby reducing power usage of the electronic device.
[0235] In some embodiments, the first input corresponds to a request to move the first object away from the individual viewpoint (1034a). In some embodiments, in response to receiving a first portion of the first input and before moving the first object away from the individual viewpoint (1034b) (e.g., in response to detecting a user's hand performing a pinch down gesture in which the tip of the index finger approaches and touches the tip of the thumb, and before the pinch hand shape is subsequently moved), in accordance with determining that the first size of the first object satisfies one or more criteria, including a criterion that is satisfied when the first size does not correspond to a current distance between the first object and the individual viewpoint (e.g., if the current size of the first object when the first portion of the first input was detected was not based on the distance between the user's viewpoint and the first object), the electronic device scales the first object to have a third size that is different from the first size, based on the current distance between the first object and the individual viewpoint (1034c), such as if the object 906a was not sized based on the current distance between the object and the viewpoint of the user 926 in FIG. 9A. Thus, in some embodiments, in response to the first portion of the first input, the electronic device appropriately sizes the first object based on a current distance between the first object and the user's viewpoint. The third size is optionally larger or smaller than the first size, depending on the current distance between the first object and the user's viewpoint. Scaling the first object upon detecting the first portion of the movement input ensures that the first object is appropriately sized for its current distance from the user's viewpoint, facilitating subsequent interaction with the first object and thereby improving user-device interaction.
[0236] It should be understood that the particular order in which the operations in method 1000 are described is merely exemplary and does not indicate that the described order is the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein.
[0237] 11A-11E illustrate examples of electronic devices that selectively resist movement of objects within a three-dimensional environment, according to some embodiments.
[0238] 11A shows electronic device 101 displaying a three-dimensional environment 1102 from the perspective of user 1126, shown in an overhead view (e.g., facing the back wall of the physical environment in which device 101 is located), via a display generating component (e.g., display generating component 120 of FIG. 1 ). As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes a display generating component (e.g., a touchscreen) and multiple image sensors (e.g., image sensor 314 of FIG. 3 ). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101 could use to capture one or more images of a user or a part of a user (e.g., one or more of the user's hands) while the user interacts with electronic device 101. In some embodiments, the user interface shown and described below may also be realized on a head-mounted display that includes display generating components that display the user interface or three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face). Device 101 optionally includes one or more buttons (e.g., physical buttons), which may optionally be a power button 1140 and volume control buttons 1141.
[0239] 11A , device 101 captures one or more images of the physical environment around device 101 (e.g., operating environment 100), including one or more objects in the physical environment around device 101. In some embodiments, device 101 displays a representation of the physical environment in three-dimensional environment 1102. For example, three-dimensional environment 1102 includes coffee table representation 1122a (corresponding to table 1122b in the overhead view), which is optionally a representation of a physical coffee table in the physical environment, and three-dimensional environment 1102 includes sofa representation 1124a (corresponding to sofa 1124b in the overhead view), which is optionally a representation of a physical sofa in the physical environment.
[0240] 11A , three-dimensional environment 1102 also includes virtual objects 1104a (corresponding to object 1104b in the overhead view), 1106a (corresponding to object 1106b in the overhead view), 1107a (corresponding to object 1107b in the overhead view), and 1109a (corresponding to object 1109b in the overhead view). Virtual objects 1104a and 1106a are optionally at a relatively small distance from the viewpoint of user 1126, and virtual objects 1107a and 1109a are optionally at a relatively large distance from the viewpoint of user 1126. In FIG. 11A , virtual object 1109a is at the farthest distance from the viewpoint of user 1126. In some embodiments, virtual object 1107a is a valid drop target for virtual object 1104a, and virtual object 1109a is an invalid drop target for virtual object 1106a. For example, virtual object 1107a is a user interface of an application (e.g., a messaging user interface) configured to accept and / or display virtual object 1104a, which is optionally a two-dimensional photograph. Virtual object 1109a is a user interface of an application (e.g., a content browsing user interface) that is not capable of accepting and / or displaying virtual object 1106a, which is optionally also a two-dimensional photograph. In some embodiments, virtual objects 1104a and 1106a are optionally one or more of a user interface of an application that contains content (e.g., a quick look window that displays a photograph), a three-dimensional object (e.g., a virtual clock, a virtual ball, a virtual car, etc.), or any other element displayed by device 101 that is not included in the physical environment of device 101.
[0241] In some embodiments, the virtual objects are displayed in the three-dimensional environment 1102 at their respective orientations relative to the viewpoint of the user 1126 (e.g., before receiving input in the three-dimensional environment 1102 to interact with the virtual objects, as described below). As shown in FIG. 11A , virtual objects 1104a and 1106a have a first orientation (e.g., the forward-facing surfaces of virtual objects 1104a and 1106a facing the viewpoint of user 1126 are tilted / slightly angled upward with respect to the viewpoint of user 1126), virtual object 1107a has a second orientation that is different from the first orientation (e.g., the forward-facing surface of virtual object 1107a facing the viewpoint of user 1126 is tilted / slightly angled left with respect to the viewpoint of user 1126, as indicated by 1107b in the overhead view), and virtual object 1109a has a third orientation that is different from the first and second orientations (e.g., the forward-facing surface of virtual object 1109a facing the viewpoint of user 1126 is tilted / slightly angled right with respect to the viewpoint of user 1126, as indicated by 1109b in the overhead view). 11A is merely exemplary, and other orientations are possible. For example, the objects optionally all share the same orientation in the three-dimensional environment 1102.
[0242] In some embodiments, a shadow of a virtual object is optionally displayed by device 101 over a valid drop target for that virtual object. For example, in Figure 11A, the shadow of virtual object 1104a is displayed overlaid on virtual object 1107a, which is a valid drop target for virtual object 1104a. In some embodiments, the relative size of the shadow of virtual object 1104a optionally changes in response to changes in the position of virtual object 1104a with respect to virtual object 1107a; thus, in some embodiments, the shadow of virtual object 1104a indicates the distance between objects 1104a and 1107a. For example, movement of virtual object 1104a toward virtual object 1107a (e.g., away from the viewpoint of user 1126) optionally decreases the size of the shadow of virtual object 1104a overlaid on virtual object 1107a, and movement of virtual object 1104a away from virtual object 1107a (e.g., closer to the viewpoint of user 1126) optionally increases the size of the shadow of virtual object 1104a overlaid on virtual object 1107a. In Figure 11A, because virtual object 1106a is an invalid drop target for virtual object 1109a, the shadow of virtual object 1109a is not displayed overlaid on virtual object 1106a.
[0243] In some embodiments, device 101 resists movement of the object along a particular path or in a particular direction within three-dimensional environment 1102. For example, device 101 resists movement of the object along a path that includes another object and / or in a direction through another object within three-dimensional environment 1102. In some embodiments, movement of the first object through another object is prevented pursuant to a determination that the other object is a valid drop target for the first object. In some embodiments, movement of the first object through another object is not resisted pursuant to a determination that the other object is an invalid drop target for the first object. Additional details regarding the above object movements are provided below with reference to method 1200.
[0244] 11A , hand 1103a (e.g., in hand state A) is providing movement input directed toward object 1104a, and hand 1105a (e.g., in hand state A) is providing movement input directed toward object 1106a. Hand 1103a is optionally providing input to move object 1104a farther from the viewpoint of user 1126 toward virtual object 1107a in three-dimensional environment 1102, and hand 1105a is optionally providing input to move object 1106a farther from the viewpoint of user 1126 toward virtual object 1109a in three-dimensional environment 1102. In some embodiments, such movement input includes moving the user's hand away from the body of user 1126 while the user's hand is in a pinch hand shape (e.g., while the tips of the thumb and index finger of the hand are touching). 11A-11B, device 101 optionally detects that hand 1103a moves away from the body of user 1126 while in the pinch hand shape, and device 101 optionally detects that hand 1105a moves away from the body of user 1126 while in the pinch hand shape. While multiple hands and corresponding inputs are shown in FIGS. 11A-11E, it should be understood that such hands and inputs need not be detected simultaneously by device 101. Rather, in some embodiments, device 101 responds independently to the hands and / or inputs shown and described in response to independently detecting such hands and / or inputs.
[0245] In response to the movement input detected in Figure 11A, device 101 moves objects 1104a and 1106a in three-dimensional environment 1102 accordingly, as shown in Figure 11B. In Figure 11A, hands 1104a and 1106a, optionally, have the same direction and magnitude of movement, as described above. In response to a given magnitude of movement of hand 1103a away from the body of user 1126, device 101 moves object 1104a away from the viewpoint of user 1126 toward virtual object 1107a, and the movement of virtual object 1104a is stopped by virtual object 1107a in three-dimensional environment 1102 (e.g., when virtual object 1104a reaches and / or collides with virtual object 1107a), as shown in the overhead view of Figure 11B. In response to the same given magnitude of movement of hand 1105a away from the body of user 1126, device 101 moves object 1104a away from the viewpoint of user 1126 by a distance greater than the distance covered by object 1106a, as shown in the overhead view of FIG. 11B. As described above, virtual object 1107a is optionally a valid drop target for virtual object 1104a, and virtual object 1109a is optionally an invalid drop target for virtual object 1106a. Thus, when virtual object 1104a is moved by device 101 in the direction of virtual object 1107a in response to a given magnitude of movement of hand 1103a, the movement of virtual object 1107a through virtual object 1104a is optionally resisted by device 101 when virtual object 1104a reaches / touches at least a portion of the surface of virtual object 1107a, as shown in the overhead view of FIG. 11B.On the other hand, since the virtual object 1106a is moved by the device 101 in the direction of the virtual object 1109a in response to a given magnitude of movement of the hand 1105a, when the virtual object 1106a reaches / touches the surface of the virtual object 1109a, the movement of the virtual object 1106a is not resisted, but in FIG. 11B, as shown in the overhead view of FIG. 11B, the virtual object 1106a has not yet reached the virtual object 1109a.
[0246] Additionally, in some embodiments, device 101 automatically adjusts the orientation of an object to correspond to another object or surface when the object approaches that other object if the other object is a valid drop target for that object. For example, depending on a given magnitude of movement of hand 1103a, when virtual object 1104a is moved within a threshold distance (e.g., 0.1, 0.5, 1, 3, 6, 12, 24, 36, or 48 cm) of the surface of virtual object 1107a, device 101 optionally adjusts the orientation of virtual object 1107a to correspond to and / or be parallel to the approached surface of virtual object 1104a (e.g., as shown in the overhead view of FIG. 11B ), since virtual object 1107a is a valid drop target for virtual object 1104a. In some embodiments, device 101 ceases automatically adjusting the orientation of an object to correspond to another object or surface when the object approaches that other object if the other object is an invalid drop target for that object. For example, when virtual object 1106a is moved within a threshold distance (e.g., 0.1, 0.5, 1, 3, 6, 12, 24, 36, or 48 cm) of the surface of virtual object 1109a in response to a given magnitude of movement of hand 1105a, device 101 ceases adjusting the orientation of virtual object 1106a to correspond to and / or be parallel to the approached surface of virtual object 1109a (e.g., as shown in the overhead view of FIG. 11B ) because virtual object 1109a is an invalid drop target for virtual object 1106a.
[0247] Additionally, in some embodiments, when an individual object is moved within a threshold distance of a surface of the object (e.g., physical or virtual), device 101 displays a badge on the individual object indicating whether the object is a valid or invalid drop target for the individual object. In FIG. 11B , object 1107a is a valid drop target for object 1104a, and therefore device 101 displays badge 1125 overlaid on the upper right corner of object 1104a indicating that object 1107a is a valid drop target for object 1104a when virtual object 1104a is moved within a threshold distance (e.g., 0.1, 0.5, 1, 3, 6, 12, 24, 36, or 48 cm) of the surface of virtual object 1107a. In some embodiments, badge 1125 optionally includes one or more symbols or characters (e.g., a “+” symbol indicating that virtual object 1107a is a valid drop target for virtual object 1104a). In another example, if virtual object 1107a is an invalid drop target for virtual object 1104a, badge 1125 optionally includes one or more symbols or characters (e.g., a "-" symbol or an "x" symbol) indicating that virtual object 1107a is an invalid drop target for virtual object 1104a. In some embodiments, when an individual object is moved within a threshold distance of the object, if the object is a valid drop target for the individual object, device 101 resizes the individual object to indicate that the object is a valid drop target for the individual object. For example, in FIG. 11B , when virtual object 1104a is moved within a threshold distance of the surface of virtual object 1107a, virtual object 1104a is scaled down (or up) in size (e.g., angularly) in three-dimensional environment 1102.The size to which object 1104a is scaled is, optionally, based on the size of object 1107a and / or the size of the area within object 1104a that can accept object 1107a (e.g., the larger object 1107a, the larger the scaled size of object 1104a). Further details of valid and invalid drop targets and associated indications displayed, and other responses of device 101, are described with reference to methods 1000, 1200, 1400, and / or 1600.
[0248] Additionally, in some embodiments, device 101 controls the size of objects included in three-dimensional environment 1102 based on the object's distance from the user's 1126 viewpoint to prevent objects from consuming a large portion of the user's 1126 field of view from their current viewpoint. Thus, in some embodiments, objects are associated with appropriate or optimal sizes for their current distance from the user's 1126 viewpoint, and device 101 automatically resizes the objects to fit those appropriate or optimal sizes. However, in some embodiments, device 101 does not adjust the size of the objects until user input to move the objects is detected. For example, in FIG. 11A , objects 1104a and 1106a are displayed by device 101 at a first size in the three-dimensional environment. In response to detecting input provided by hand 1103a to move object 1104a within three-dimensional environment 1102 and hand 1105a to move object 1106a within three-dimensional environment 1102, device 101 optionally increases the size of objects 1104a and 1106a within three-dimensional environment 1102, as shown in the overhead view of Figure 11B. The increased size of objects 1104a and 1106a optionally corresponds to a current distance of objects 1104a and 1106a from a viewpoint of user 1126. Additional details regarding controlling the size of objects based on the distance of the objects from a viewpoint of a user are described with reference to the Figure 9 series of diagrams and method 1000.
[0249] In some embodiments, device 101 applies varying levels of resistance to movement of a first object along the surface of a second object depending on whether the second object is a valid drop target for the first object. For example, in Figure 11B, hand 1103b (e.g., in hand state B) is providing an upward diagonal movement input directed toward object 1104a, and hand 1105b (e.g., in hand state B) is providing an upward diagonal movement input to object 1106a while object 1104a is already in contact with object 1107a. In hand state B (e.g., while the hand is in a pinch hand shape (e.g., while the tips of the thumb and index finger of the hand are touching)), hand 1103b is optionally providing input to move object 1104a diagonally (e.g., across to the right) further from the perspective of user 1126 into the surface of virtual object 1107a in three-dimensional environment 1102, and hand 1105b is optionally providing input to move object 1106a diagonally (e.g., across to the right) further from the perspective of user 1126 into the surface of virtual object 1109a in three-dimensional environment 1102.
[0250] In some embodiments, in response to a given amount of hand movement, device 101 moves a first object by different amounts within three-dimensional environment 1102 depending on whether the first object is in contact with the surface of a second object and whether the second object is a valid drop target for the first object. For example, in Figure 11B, the movement amounts of hands 1103b and 1105b are optionally the same. In response, as shown in Figure 11C, device 101 moves object 1106a diagonally in three-dimensional environment 1102 rather than moving object 1104a laterally and / or away from the viewpoint of user 1126 in three-dimensional environment 1102. 11C , in response to a diagonal movement of hand 1103b within three-dimensional environment 1102, optionally including a right-lateral component and a component away from the viewpoint of user 1126, device 101 resisted (e.g., did not allow) the movement of object 1104a away from the viewpoint of user 1126 (e.g., in accordance with the component of the movement of hand 1303b away from the viewpoint of user 1126) because object 1104a was in contact with object 1107a, object 1107a was a valid drop target for object 1104a, and the component of the movement of hand 1303b away from the viewpoint of user 1126 was not sufficient to break through object 1107a, as described below. 11C, device 101 is moving object 1104a a relatively small amount (e.g., less than the lateral movement of object 1106a) across the surface of object 1107a (e.g., in accordance with the right lateral component of the movement of hand 1303b) because object 1104a is in contact with object 1107a and object 1107a is a valid drop target for object 1104a. Additionally or alternatively, in FIG. 11C, in response to the movement of hand 1105b, device 101 is moving virtual object 1106a diagonally within three-dimensional environment 1102, such that virtual object 1106a is displayed by device 101 behind virtual objects 1109a and 1107a from the perspective of user 1126, as shown in the overhead view of FIG. 11C.Because virtual object 1109a is an invalid drop target for virtual object 1106a, movement of virtual object 1106a diagonally through virtual object 1109a is optionally not resisted by device 101, and the lateral movement of object 1106a and the movement of object 1106a away from the viewpoint of user 1126 (e.g., according to the right lateral component of the movement of hand 1105b and according to the component of the movement of hand 1303b away from the viewpoint of user 1126, respectively) is greater than the lateral movement of object 1104a and the movement of object 1104a away from the viewpoint of user 1126.
[0251] In some embodiments, device 101 requires at least a threshold magnitude of movement of an individual object through an object to allow the individual object to pass through the object when the individual object is in contact with the object's surface. For example, in Figure 11C, hand 1103c (e.g., in hand state C) is providing movement input directed at virtual object 1104a to move virtual object 1104a through the surface of virtual object 1107a from the perspective of user 1126. In hand state C (e.g., while the hand is in a pinch hand shape (e.g., while the tips of the thumb and index finger of the hand are touching)), hand 1103c is optionally providing input to move object 1104a further into (e.g., vertically) the surface of virtual object 1107a in three-dimensional environment 1102 from the perspective of user 1126. In response to the first portion of the movement input moving virtual object 1104a into / through virtual object 1107a, device 101 optionally resists the movement of virtual object 1104a through virtual object 1107a. When hand 1103c applies a larger magnitude of movement in a second portion of the movement input moving virtual object 1104a through virtual object 1107a, device 101 optionally resists the movement with an increasing resistance level, optionally proportional to the increasing magnitude of the movement. In some embodiments, when the magnitude of the movement directed at virtual object 1104a reaches and / or exceeds a discrete magnitude threshold (e.g., corresponding to a movement of 0.3, 0.5, 1, 2, 3, 5, 10, 20, 40, or 50 cm), device 101 moves virtual object 1104a through virtual object 1107a, as shown in FIG. 11D .
[0252] 11D , in response to detecting that the movement input directed at virtual object 1104a exceeds an individual magnitude threshold, device 101 ceases resisting the movement of virtual object 1104a through virtual object 1107a, allowing virtual object 1107a to pass through virtual object 1104a. In Figure 11D , virtual object 1104a is moved by device 101 to a location behind virtual object 1107a in three-dimensional environment 1102, as shown in the overhead view of Figure 11D . In some embodiments, upon detecting that the movement input directed at virtual object 1104a exceeds an individual magnitude threshold, device 101 provides a visual indication 1116 in three-dimensional environment 1102 (e.g., on the surface of object 1107a at the location through which object 1104a passed) that indicates that virtual object 1104a has moved through the surface of virtual object 1107a from the perspective of user 1126. For example, in FIG. 11D, device 101 displays ripples 1116 on the surface of virtual object 1107a from the perspective of user 1126, indicating that virtual object 1104a has moved through virtual object 1107a in accordance with the movement input.
[0253] In some embodiments, device 101 provides a visual indication of the presence of virtual object 1104a behind virtual object 1107a in three-dimensional environment 1102. For example, in Figure 11D, device 101 alters the appearance of virtual object 1104a and / or the appearance of virtual object 1107a so that the distinct location of virtual object 1104a is discernible from the perspective of user 1126, even though virtual object 1104a is located behind virtual object 1107a in three-dimensional environment 1102. In some embodiments, the visual indication of virtual object 1104a behind virtual object 1107a in three-dimensional environment 1102 is a faded or ghosted version of object 1104a displayed through (e.g., overlaid on) object 1107a, an outline of object 1104a displayed through (e.g., overlaid on) object 1107a, etc. In some embodiments, the device 101 increases the transparency of the virtual object 1107a to provide a visual indication of the presence of the virtual object 1104a behind the virtual object 1107a in the three-dimensional environment 1102.
[0254] In some embodiments, lateral movement of an individual object, or movement of an individual object further from the viewpoint of user 1126, while the individual object is behind an object in three-dimensional environment 1102 is not resisted by device 101. For example, in Figure 11E, when hand 1103d provides a movement input directed at virtual object 1104a to move virtual object 1104a laterally to the right to a new location behind virtual object 1107a in three-dimensional environment 1102, device 101 optionally moves object 1104a to the new location in accordance with the movement input without resisting the movement. Further, in some embodiments, device 101 will update the display of the visual indication of object 1104a in three-dimensional environment 1102 (e.g., a ghost, outline, etc. of virtual object 1104a) to have a different size and / or a new portion through the surface of virtual object 1107a that corresponds to the new location of object 1104a behind virtual object 1107a from the perspective of user 1126 based on the updated distance of object 1104a from the perspective of user 1126.
[0255] In some embodiments, movement of virtual object 1104a from behind virtual object 1107a to a discrete location in front of virtual object 1107a (e.g., through virtual object 1107a) from the perspective of user 1126 is not resisted by device 101. In Figure 11D, hand 1103d (e.g., in hand state D) is providing movement input directed at virtual object 1104a to move virtual object 1104a from behind virtual object 1107a to a discrete location in front of virtual object 1107a, from the perspective of user 1126, along a path through virtual object 1107a, as shown in Figure 11E. In hand state D (e.g., while the hand is in a pinch hand shape (e.g., while the tips of the thumb and index finger of the hand are touching)), hand 1103d is optionally providing input to move object 1104a closer to the viewpoint of user 1126 and within (e.g., vertically) the back of virtual object 1107a in three-dimensional environment 1102.
[0256] In Figure 11E, in response to detecting movement of virtual object 1104a from behind virtual object 1107a to in front of virtual object 1107a, device 101 moves virtual object 1104a through virtual object 1107a to a discrete location in front of virtual object 1107a within three dimensional environment 1102 from the perspective of user 1126, as shown in the overhead view of Figure 11E. In Figure 11E, movement of virtual object 1104a through virtual object 1107a is not resisted by device 101 as virtual object 1104a is moved to a discrete location within three dimensional environment 1102. It should be understood that in some embodiments, subsequent movement of virtual object 1104a (e.g., in respons...
Claims
1. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying, via the display generation component, a first object within a three-dimensional environment, and detecting, while the first object is selected for movement within the three-dimensional environment, via the one or more input devices, a first input corresponding to movement of a discrete body part of a user of the electronic device within a physical environment in which the display generation component is located; In response to detecting the first input, moving the first object in a first output direction within the three-dimensional environment in accordance with the movement of the individual part of the body of the user in the physical environment in the first input direction in accordance with determining that the first input includes movement of the individual part of the body of the user in the physical environment in the first input direction, wherein the movement of the first object in the first output direction has a first relationship to the movement of the individual part of the body of the user in the physical environment in the first input direction; in accordance with determining that the first input includes movement of the individual part of the body of the user within the physical environment in a second input direction different from the first input direction, moving the first object within the three dimensional environment in a second output direction different from the first output direction in accordance with the movement of the individual part of the body of the user within the physical environment in the second input direction, wherein the movement of the first object in the second output direction has a second relationship to the movement of the individual part of the body of the user within the physical environment in the second input direction that is different from the first relationship.
2. a magnitude of the movement of the first object in the first output direction is independent of a speed of the movement of the individual part of the body of the user in the first input direction; 2. The method of claim 1, wherein the magnitude of the movement of the first object in the second output direction is independent of the speed of the movement of the individual part of the body of the user in the second input direction.
3. the first output direction is horizontal with respect to a viewpoint of the user within the three-dimensional environment; The method of claim 1 or 2, wherein the second output direction is perpendicular to the viewpoint of the user within the three-dimensional environment.
4. the movement of the discrete part of the body of the user within the physical environment in the first input direction and the second input direction has a first magnitude; the movement of the first object in the first output direction has a second magnitude greater than the first magnitude; The method of claim 3 , wherein the movement of the first object in the second output direction has a third magnitude that is greater than the first magnitude and different from the second magnitude.
5. 5. The method of claim 1, wherein the first relationship is based on an offset between a second individual part of the body of the user and the individual part of the body of the user, and the second relationship is based on the offset between the second individual part of the body of the user and the individual part of the body of the user.
6. the first output direction corresponds to movement away from the user's viewpoint within the three-dimensional environment; the second output direction corresponds to movement of the user within the three-dimensional environment toward the viewpoint; the movement of the first object in the first output direction is increased by the movement of the discrete portion of the user in the first input direction by a first value based on a distance between the portion of the user and a location corresponding to the first object; 6. The method of claim 1, wherein the movement of the first object in the second output direction is increased by a second value different from the first value based on a distance between a viewpoint of the user and the location corresponding to the first object in the three-dimensional environment.
7. 7. The method of claim 6, wherein the first value changes as the movement of the discrete part of the user in the first input direction progresses and / or the second value changes as the movement of the discrete part of the user in the second input direction progresses.
8. 8. The method of claim 7, wherein the first value changes in a first manner as the movement of the discrete part of the user in the first input direction progresses, and the second value changes in a second manner different from the first manner as the movement of the discrete part of the user in the second input direction progresses.
9. 9. The method of claim 8, wherein the first value remains constant during a given portion of the movement of the discrete part of the user in the first input direction, and the second value does not remain constant during a given portion of the movement of the discrete part of the user in the second input direction.
10. The method of claim 6 , wherein the first and second multipliers are based on a ratio of distance to the user's arm length.
11. The method of claim 6 , wherein the second value is based on the position of the discrete part of the user when the first object is selected for movement.
12. detecting, via the one or more input devices, a respective movement of the respective portion of the user in a direction horizontal to a viewpoint of the user within the three-dimensional environment while the first object is selected for movement within the three-dimensional environment; 12. The method of claim 1, further comprising: in response to detecting the individual movements of the individual parts of the user, updating a location of the first object in the three-dimensional environment based on noise-reduced individual movements of the individual parts of the user.
13. When the individual movement of the individual portion of the user is detected, in accordance with a determination that a location corresponding to the first object is at a first distance from the individual portion of the user, the individual movement of the individual portion of the user is adjusted based on a first amount of noise reduction to generate an adjusted movement used to update the location of the first object within the three-dimensional environment; 13. The method of claim 12, wherein, when the individual movement of the individual portion of the user is detected, in accordance with a determination that the location corresponding to the first object is at a second distance from the individual portion of the user that is less than the first distance, the individual movement of the individual portion of the user is adjusted based on a second amount that is less than the first amount of noise reduction used to generate an adjusted movement used to update the location of the first object within the three-dimensional environment.
14. 14. The method of claim 1, further comprising, while the first object is selected for movement and during the first input, controlling an orientation of the first object within the three-dimensional environment in a plurality of directions according to corresponding orientation control portions of the first input.
15. detecting when the first object is within a threshold distance of a surface within the three-dimensional environment while controlling the orientation of the first object in the plurality of directions; 15. The method of claim 14, further comprising: in response to detecting the first object being within the threshold distance of the surface in the three-dimensional environment, updating one or more orientations of the first object within the three-dimensional environment based on an orientation of the surface.
16. detecting that the first object is no longer selected for movement while controlling the orientation of the first object in the plurality of directions; 16. The method of claim 14 or 15, further comprising: in response to detecting that the first object is no longer selected for movement, updating one or more orientations of the first object within the three-dimensional environment to be based on a default orientation of the first object within the three-dimensional environment.
17. The first input is made while the discrete portion of the user is within a threshold distance of a location corresponding to the first object, and the method further comprises: while the first object is selected for movement, and a second input corresponding to movement of the first object within the three-dimensional environment, wherein during the second input, the discrete portion of the user is further than the threshold distance from the location corresponding to the first object; according to a determination that the first object is a two-dimensional object, moving the first object within the three-dimensional environment in accordance with the second input while an orientation of the first object relative to a viewpoint of the user within the three-dimensional environment remains constant; 17. The method of claim 14, further comprising: in accordance with a determination that the first object is a three-dimensional object, moving the first object within the three-dimensional environment in accordance with the second input while an orientation of the first object relative to a surface within the three-dimensional environment remains constant.
18. While the first object is selected for movement and during the first input, moving the first object within the three-dimensional environment in accordance with the first input while maintaining an orientation of the first object relative to a viewpoint of the user within the three-dimensional environment; detecting that the first object is within a threshold distance of a second object in the three-dimensional environment after moving the first object while maintaining the orientation of the first object relative to the viewpoint of the user; 18. The method of claim 1, further comprising: in response to detecting that the first object is within the threshold distance of the second object, updating the orientation of the first object in the three-dimensional environment based on an orientation of the second object, independent of an orientation of the first object relative to the viewpoint of the user.
19. Before the first object is selected for movement within the three-dimensional environment, the first object has a first size within the three-dimensional environment, and the method further comprises:
19. The method of claim 1, further comprising, in response to detecting a selection of the first object for movement within the three-dimensional environment, scaling the first object to have a second size within the three-dimensional environment that is different from the first size, the second size being based on a distance between a location corresponding to the first object in the three-dimensional environment at the time the selection of the first object for movement was detected and a viewpoint of the user.
20. 20. The method of claim 1, wherein the first object is selected for movement within the three-dimensional environment in response to detecting a second input that includes the discrete portion of the user performing a first gesture while the user's gaze is directed toward the first object, the first gesture continuing to maintain a first shape for a threshold time period.
21. 20. The method of claim 1, wherein the first object is selected for movement within the three-dimensional environment in response to detecting a second input that includes movement of a discrete portion of the user toward a viewpoint of the user within the three-dimensional environment that is greater than a movement threshold while the user's gaze is directed toward the first object.
22. While the first object is selected for movement and during the first input, Detecting that the first object is within a threshold distance of a second object in the three-dimensional environment; In response to detecting the first object being within the threshold distance of the second object, displaying, via the display generating component, a first visual indication that the second object is a valid drop target for the first object in accordance with determining that the second object is a valid drop target for the first object; 22. The method of claim 1, further comprising: displaying, via the display generating component, a second visual indication indicating that the second object is not a valid drop target for the first object in accordance with a determination that the second object is not a valid drop target for the first object.
23. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions displaying, via a display generation component, a first object within a three-dimensional environment, and detecting, while the first object is selected for movement within the three-dimensional environment, via one or more input devices, a first input corresponding to movement of a discrete body part of a user of the electronic device within a physical environment in which the display generation component is located; In response to detecting the first input, moving the first object in a first output direction within the three-dimensional environment in accordance with the movement of the individual part of the body of the user in the physical environment in the first input direction in accordance with determining that the first input includes movement of the individual part of the body of the user in the physical environment in the first input direction, wherein the movement of the first object in the first output direction has a first relationship to the movement of the individual part of the body of the user in the physical environment in the first input direction; In accordance with determining that the first input includes movement of the individual part of the body of the user within the physical environment in a second input direction different from the first input direction, moving the first object within the three-dimensional environment in a second output direction different from the first output direction in accordance with the movement of the individual part of the body of the user within the physical environment in the second input direction, wherein the movement of the first object in the second output direction moves the first object in the second output direction having a second relationship different from the first relationship to the movement of the individual part of the body of the user within the physical environment in the second input direction.
24. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: displaying, via a display generation component, a first object within a three-dimensional environment, and detecting, while the first object is selected for movement within the three-dimensional environment, via one or more input devices, a first input corresponding to movement of a discrete body part of a user of the electronic device within a physical environment in which the display generation component is located; In response to detecting the first input, moving the first object in a first output direction within the three-dimensional environment in accordance with the movement of the individual part of the body of the user in the physical environment in the first input direction in accordance with determining that the first input includes movement of the individual part of the body of the user in the physical environment in the first input direction, wherein the movement of the first object in the first output direction has a first relationship to the movement of the individual part of the body of the user in the physical environment in the first input direction; a first output direction that is different from the first relationship to the movement of the individual part of the body of the user in the physical environment in the second input direction;
25. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a first object within a three-dimensional environment, and detecting, while the first object is selected for movement within the three-dimensional environment, via one or more input devices, a first input corresponding to movement of a discrete body part of a user of the electronic device within a physical environment in which the display generation component is located; In response to detecting the first input, moving the first object in a first output direction within the three-dimensional environment in accordance with the movement of the individual part of the body of the user in the physical environment in the first input direction in accordance with determining that the first input includes movement of the individual part of the body of the user in the physical environment in the first input direction, wherein the movement of the first object in the first output direction has a first relationship to the movement of the individual part of the body of the user in the physical environment in the first input direction; and means for, in accordance with determining that the first input includes movement of the individual part of the body of the user within the physical environment in a second input direction different from the first input direction, moving the first object within the three dimensional environment in a second output direction different from the first output direction in accordance with the movement of the individual part of the body of the user within the physical environment in the second input direction, wherein the movement of the first object in the second output direction has a second relationship to the movement of the individual part of the body of the user within the physical environment in the second input direction that is different from the first relationship.
26. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a first object within a three-dimensional environment, and detecting, while the first object is selected for movement within the three-dimensional environment, via one or more input devices, a first input corresponding to movement of a discrete body part of a user of the electronic device within a physical environment in which the display generation component is located; In response to detecting the first input, moving the first object in a first output direction within the three-dimensional environment in accordance with the movement of the individual part of the body of the user in the physical environment in the first input direction in accordance with determining that the first input includes movement of the individual part of the body of the user in the physical environment in the first input direction, wherein the movement of the first object in the first output direction has a first relationship to the movement of the individual part of the body of the user in the physical environment in the first input direction; and means for, in accordance with a determination that the first input includes movement of the individual part of the body of the user within the physical environment in a second input direction different from the first input direction, moving the first object within the three-dimensional environment in a second output direction different from the first output direction in accordance with the movement of the individual part of the body of the user within the physical environment in the second input direction, wherein the movement of the first object in the second output direction has a second relationship to the movement of the individual part of the body of the user within the physical environment in the second input direction, the second relationship being different from the first relationship.
27. 23. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 1 to 22.
28. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method according to any one of claims 1 to 22.
29. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising:
23. An information processing device comprising means for carrying out the method according to any one of claims 1 to 22.
30. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying, via the display generation component, the three-dimensional environment including a first object at a first location within the three-dimensional environment, the first object having a first size within the three-dimensional environment and occupying a first amount of a field of view from a distinct viewpoint; receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, a first input via the one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, upon determining that the first input corresponds to a request to move the first object away from the individual viewpoint; moving the first object away from the distinct viewpoint from the first location to a second location within the three-dimensional environment in accordance with the first input, the second location being farther from the distinct viewpoint than the first location; scaling the first object so that when the first object is located at the second location, the first object has a second size in the three-dimensional environment that is larger than the first size and occupies a second amount of the field of view from the distinct viewpoint, the second amount being smaller than the first amount.
31. 31. The method of claim 30, further comprising, while receiving the first input, in accordance with the determination that the first input corresponds to the request to move the first object away from the distinct viewpoint, continuously scaling the first object to increasing sizes as the first object moves further from the distinct viewpoint.
32. The first object is an object of a first type, and the three-dimensional environment further includes a second object, the second object being an object of a second type different from the first type, and the method further comprises: receiving, while displaying the three-dimensional environment including the second object at a third location within the three-dimensional environment, the second object having a third size within the three-dimensional environment and occupying a third amount of the field of view from the distinct viewpoint, a second input via the one or more input devices corresponding to a request to move the second object away from the third location within the three-dimensional environment; in response to receiving the third input and in accordance with a determination that the second input corresponds to a request to move the second object away from the individual viewpoint; 32. The method of claim 30 or 31, further comprising: moving the second object from the third location to a fourth location within the three-dimensional environment, away from the individual viewpoint, in accordance with the second input, wherein the fourth location is farther from the individual viewpoint than the third location, without scaling the second object, such that when the second object is located at the fourth location, the second object has the third size within the three-dimensional environment and occupies a fourth amount of the field of view from the individual viewpoint that is less than the third amount.
33. the second object is displayed with a control user interface for controlling one or more operations associated with the second object; when the second object is displayed at the third location, the control user interface is displayed at the third location and has a fourth size within the three-dimensional environment; 33. The method of claim 32, wherein when the second object is displayed at the fourth location, the control user interface is displayed at the fourth location and has a fifth size in the three-dimensional environment that is larger than the fourth size.
34. While displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, the first object has the first size within the three-dimensional environment and the distinct viewpoint is a first viewpoint, detecting a movement of the user's viewpoint from the first viewpoint to a second viewpoint that changes a distance between the user's viewpoint and the first object; 34. The method of claim 30, further comprising: in response to detecting the movement of the viewpoint from the first viewpoint to the second viewpoint, updating the display of the three-dimensional environment to be from the second viewpoint without scaling a size of the first object at the first location in the three-dimensional environment.
35. The first object is an object of a first type, and the three-dimensional environment further includes a second object, the second object being an object of a second type different from the first type, and the method further comprises: while displaying the three-dimensional environment including the second object at a third location within the three-dimensional environment, the second object having a third size within the three-dimensional environment, the viewpoint of the user being the first viewpoint, and detecting a movement of the viewpoint from the first viewpoint to the second viewpoint that changes a distance between the viewpoint of the user and the second object; In response to detecting the movement of the individual viewpoint, updating a representation of the three-dimensional environment from the second perspective; 35. The method of claim 34, further comprising scaling a size of the second object at the third location to a fourth size within the three-dimensional environment that is different from the third size.
36. detecting, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, a movement of the user's viewpoint within the three-dimensional environment from a first viewpoint to a second viewpoint that changes a distance between the viewpoint and the first object; In response to detecting the movement of the viewpoint, updating the display of the three-dimensional environment to be from the second viewpoint without scaling a size of the first object at the first location within the three-dimensional environment; While displaying the first object at the first location within the three-dimensional environment from the second viewpoint, receiving, via the one or more input devices, a second input corresponding to a request to move the first object away from the first location within the three-dimensional environment to a third location within the three-dimensional environment that is farther from the second distinct location than the first location; 36. The method of claim 30, further comprising: scaling a size of the first object to a third size different from the first size based on a distance between the first object and the second viewpoint when a start of the second input is detected while detecting the second input and before moving the first object away from the first location.
37. The three-dimensional environment further includes a second object at a third location within the three-dimensional environment, and the method further comprises: In response to receiving the first input, displaying the first object at a fourth location within the three-dimensional environment, the first object having a third size within the three-dimensional environment, in response to a determination that the first input corresponds to a request to move the first object to a fourth location within the three-dimensional environment, the fourth location being a first distance from the distinct viewpoint; 37. The method of claim 36, further comprising: displaying the first object at a third location in the three-dimensional environment, the first object having a fourth size in the three-dimensional environment different from the third size, in accordance with a determination that the first input satisfies one or more criteria including individual criteria that are satisfied when the first input corresponds to a request to move the first object to the third location in the three-dimensional environment, the third location being at the first distance from the individual viewpoint.
38. 38. The method of claim 37, wherein the fourth size of the first object is based on the size of the second object.
39. receiving, while the first object is at the third location within the three-dimensional environment and has the fourth size based on the size of the second object, a second input via the one or more input devices corresponding to a request to move the first object away from the third location within the three-dimensional environment; 39. The method of claim 38, further comprising: in response to receiving the second input, displaying the first object at a fifth size, the fifth size being independent of the size of the second object.
40. 40. The method of any one of claims 37 to 39, wherein the individual criteria are satisfied when the first input corresponds to a request to move the first object to any location within a volume in the three-dimensional environment that includes the third location.
41. 41. The method of any one of claims 37 to 40, further comprising, while receiving the first input, in accordance with a determination that the first object has moved to the third location in accordance with the first input and in accordance with a determination that the one or more criteria are satisfied, modifying an appearance of the first object to indicate that the second object is a valid drop target for the first object.
42. The one or more criteria include a criterion that is satisfied when the second object is a valid drop target for the first object and is not satisfied when the second object is not a valid drop target for the first object, and the method further comprises: In response to receiving the first input, 42. The method of any one of claims 37 to 41, further comprising: displaying the first object at the fourth location within the three-dimensional environment, the first object having the third size within the three-dimensional environment, in accordance with a determination that the individual criteria are met but the first input does not satisfy the one or more criteria because the second object is not a valid drop target for the first object.
43. In response to receiving the first input, 43. The method of claim 37, further comprising: updating an orientation of the first object relative to the individual viewpoint based on an orientation of the second object relative to the individual viewpoint in accordance with the determination that the first input satisfies the one or more criteria.
44. The three-dimensional environment further includes a second object at a third location within the three-dimensional environment, and the method further comprises: While receiving the first input, in response to determining that the first input corresponds to a request to move the first object through the third location farther from the individual viewpoint than the third location; moving the first object from the first location to the third location away from the distinct viewpoint in accordance with the first input while scaling the first object within the three-dimensional environment based on a distance between the distinct viewpoint and the first object; 44. The method of any one of claims 30 to 43, further comprising, after the first object reaches the third location, maintaining a display of the first object at the third location without scaling the first object while continuing to receive the first input.
45. Scaling the first object is pursuant to determining that the second amount of the field of view from the distinct viewpoint occupied by the first object at the second size is greater than a threshold amount of the field of view, and the method further comprises: receiving, while displaying the first object at a discrete size within the three-dimensional environment, the first object occupying a first discrete amount of the field of view from the discrete viewpoint, a second input via the one or more input devices corresponding to a request to move the first object away from the discrete viewpoint; In response to receiving the second input, 44. The method of any one of claims 30 to 43, further comprising: in accordance with a determination that the first individual amount of field of view from the individual viewpoint is less than the threshold amount of field of view, moving the first object away from the individual viewpoint in accordance with the second input without scaling a size of the first object within the three-dimensional environment.
46. The first input corresponds to the request to move the first object away from the individual viewpoint, and the method further comprises: before moving the first object away from the distinct viewpoint in response to receiving a first portion of the first input; 46. The method of any one of claims 30 to 45, further comprising: scaling the first object to have a third size different from the first size, the third size being based on the current distance between the first object and the individual viewpoint, in accordance with a determination that the first size of the first object satisfies one or more criteria, including a criterion that is met when the first size does not correspond to the current distance between the first object and the individual viewpoint.
47. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment, the first object having a first size within the three-dimensional environment and occupying a first amount of a field of view from a distinct viewpoint; receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, upon determining that the first input corresponds to a request to move the first object away from the individual viewpoint; moving the first object away from the distinct viewpoint from the first location to a second location within the three-dimensional environment in accordance with the first input, the second location being farther from the distinct viewpoint than the first location; and scaling the first object so that when the first object is located at the second location, the first object has a second size in the three-dimensional environment that is larger than the first size and occupies a second amount of the field of view from the distinct viewpoint, the second amount being smaller than the first amount.
48. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment, the first object having a first size within the three-dimensional environment and occupying a first amount of a field of view from a distinct viewpoint; receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, upon determining that the first input corresponds to a request to move the first object away from the individual viewpoint; moving the first object away from the distinct viewpoint from the first location to a second location within the three-dimensional environment in accordance with the first input, the second location being farther from the distinct viewpoint than the first location; scaling the first object such that, when the first object is located at the second location, the first object has a second size in the three-dimensional environment that is larger than the first size and occupies a second amount of the field of view from the distinct viewpoint, the second amount being smaller than the first amount.
49. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment, the first object having a first size within the three-dimensional environment and occupying a first amount of a field of view from a distinct viewpoint; means for receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, upon determining that the first input corresponds to a request to move the first object away from the individual viewpoint; moving the first object away from the distinct viewpoint from the first location to a second location within the three-dimensional environment in accordance with the first input, the second location being farther from the distinct viewpoint than the first location; means for scaling the first object such that, when the first object is located at the second location, the first object has a second size in the three-dimensional environment that is larger than the first size and occupies a second amount of the field of view from the distinct viewpoint, the second amount being smaller than the first amount.
50. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment, the first object having a first size within the three-dimensional environment and occupying a first amount of a field of view from a distinct viewpoint; means for receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, upon determining that the first input corresponds to a request to move the first object away from the individual viewpoint; moving the first object away from the distinct viewpoint from the first location to a second location within the three-dimensional environment in accordance with the first input, the second location being farther from the distinct viewpoint than the first location; means for scaling the first object so that, when the first object is located at the second location, the first object has a second size in the three-dimensional environment that is larger than the first size and occupies a second amount of the field of view from the distinct viewpoint, the second amount being smaller than the first amount.
51. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 30 to 46.
52. 47. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 30 to 46.
53. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 30 to 46.
54. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising:
47. An information processing device comprising means for carrying out the method of any one of claims 30 to 46.
55. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying, via the display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment a first distance away from the first object within the three-dimensional environment; receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, via the one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment by a second distance, the second distance being greater than the first distance; In response to receiving the first input, in accordance with a determination that the first input satisfies a first set of one or more criteria, the first set of criteria including a requirement that the first input correspond to movement through the second location within the three-dimensional environment, moving the first object the first distance away from the first location within the three-dimensional environment in accordance with the first input; in response to a determination that the first input does not satisfy the first set of one or more criteria because the first input does not correspond to movement through the second location within the three-dimensional environment; and moving the first object the second distance away from the first location within the three-dimensional environment in accordance with the first input.
56. after moving the first object the first distance from the first location within the three-dimensional environment in accordance with the first input such that the first input satisfies the first set of one or more criteria; receiving a second input via the one or more input devices corresponding to a request to move the first object a third distance away from the second location within the three-dimensional environment; In response to receiving the second input, in accordance with a determination that the second input satisfies a second set of one or more criteria, the second set of one or more criteria including a requirement that the second input correspond to a movement greater than a movement threshold, moving the first object through the second object in accordance with the second input to a third location within the three-dimensional environment; 56. The method of claim 55, further comprising: maintaining the first object at the first distance away from the first location within the three-dimensional environment in accordance with a determination that the second input does not satisfy the second set of criteria because the second input does not correspond to movement greater than the movement threshold.
57. Moving the first object through the second object to the third location within the three-dimensional environment in accordance with the second input includes:
57. The method of claim 56, comprising displaying visual feedback on a portion of the second object corresponding to the location of the first object when the first object is moved through the second object to the third location in the three-dimensional environment in accordance with the second input.
58. after moving the first object to the third location in the three-dimensional environment through the second object in accordance with the second input, the second object being between the third location and a viewpoint of the three-dimensional environment displayed via the display generation component; 58. The method of claim 56 or 57, further comprising displaying, via the display generation component, a visual indication of the first object through the second object.
59. after moving the first object to the third location in the three-dimensional environment through the second object in accordance with the second input, the second object being between the third location and a viewpoint of the three-dimensional environment displayed via the display generation component; receiving a third input via the one or more input devices corresponding to a request to move the first object while the second object remains between the first object and the viewpoint of the three-dimensional environment; 59. The method of any one of claims 56 to 58, further comprising, in response to receiving the third input, moving the first object within the three-dimensional environment in accordance with the third input.
60. Moving the first object the first distance away from the first location in the three dimensional environment in accordance with the first input includes:
60. The method of any one of claims 55 to 59, comprising displaying, via the display generation component, a visual indication that the second object is the valid drop target for the first object in accordance with a determination that a second set of one or more criteria is met, the criteria including criteria that are met when the second object is a valid drop target for the first object and criteria that are met when the first object is within a threshold distance of the second object.
61. 61. The method of claim 60, wherein displaying, via the display generating component, the visual indication that the second object is the valid drop target for the first object includes resizing the first object within the three-dimensional environment.
62. 62. The method of claim 60 or 61, wherein displaying, via the display generating component, the visual indication that the second object is the valid drop target for the first object comprises displaying, via the display generating component, a first visual indicator overlaid on the first object.
63. Moving the first object the first distance away from the first location within the three dimensional environment in accordance with the first input satisfying the first set of one or more criteria includes: moving the first object a first amount in a respective direction different from a direction through the second location in response to receiving a first portion of the first input corresponding to a first magnitude of movement in the respective direction before the first object reaches the second location; 63. The method of any one of claims 55 to 62, comprising: after the first object reaches the second location, moving the first object in the respective direction by a second amount less than the first amount in response to receiving a second portion of the first input corresponding to the first magnitude of movement in the respective direction.
64. a discrete input corresponding to movement through the second location beyond movement to the second location is directed at the first object; In accordance with determining that the discrete input has a second magnitude, the second amount of movement of the first object in the discrete direction is a first discrete amount; 64. The method of claim 63, wherein, in accordance with a determination that the discrete input has a third magnitude greater than the second magnitude, the second amount of movement of the first object in the discrete direction is a second discrete amount less than the first discrete amount.
65. 65. The method of any one of claims 55 to 64, further comprising displaying, via the display generation component, a virtual shadow of the first object overlaid on the second object while moving the first object the first distance away from the first location in the three-dimensional environment in accordance with the first input so that the first input satisfies the first set of one or more criteria, wherein as the first object is moved the first distance away from the first location in the three-dimensional environment, a size of the virtual shadow of the first object overlaid on the second object is scaled according to a change in distance between the first object and the second object.
66. 65. The method of any one of claims 55 to 64, wherein the first object is a two-dimensional object and the first distance corresponds to a distance between a point on a plane of the first object and the second object.
67. 65. The method of any one of claims 55 to 64, wherein the first object is a three-dimensional object and the first distance corresponds to the distance between a point on a surface of the first object that is closest to the second object and the second object.
68. 68. The method of any one of claims 55 to 67, wherein the first set of criteria includes a requirement that at least a portion of the first object matches at least a portion of the second object when the first object is at the second location.
69. 68. A method according to any one of claims 55 to 67, wherein the first set of criteria includes a requirement that the second object be a valid drop target for the first object.
70. the first object has a first orientation within the three-dimensional environment prior to receiving the first input, and the second object has a second orientation within the three-dimensional environment that is different from the first orientation, and the method further comprises: without receiving an orientation adjustment input to adjust an orientation of the first object to correspond to the second orientation of the second object.
70. The method of any one of claims 55 to 69, further comprising adjusting the orientation of the first object to correspond to the second orientation of the second object after moving the first object the first distance away from the first location in the three-dimensional environment in accordance with the first input so that the first input satisfies a first set of the one or more criteria.
71. the three-dimensional environment includes a third object at a fourth location within the three-dimensional environment, the second object being between the fourth location and the viewpoint of the three-dimensional environment displayed via the display generation component, and the method further comprising: receiving, while displaying the three-dimensional environment including the third object at the fourth location within the three-dimensional environment and the second object at the second location between the fourth location and the viewpoint of the three-dimensional environment, a fourth input via the one or more input devices corresponding to a request to move the third object a discrete distance through the second object to a discrete location between the second location and the viewpoint of the three-dimensional environment; 71. The method of any one of claims 55 to 70, further comprising: in response to receiving the fourth input, moving the third object the discrete distance through the second object in accordance with the fourth input to the discrete location between the second location and the viewpoint of the three-dimensional environment.
72. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment a first distance away from the first object within the three-dimensional environment; while displaying the three-dimensional environment, the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, receiving, via one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment a second distance, the second distance being greater than the first distance; In response to receiving the first input, in accordance with a determination that the first input satisfies a first set of one or more criteria, the first set of criteria including a requirement that the first input correspond to movement through the second location within the three-dimensional environment, moving the first object the first distance away from the first location within the three-dimensional environment in accordance with the first input; in response to a determination that the first input does not satisfy the first set of one or more criteria because the first input does not correspond to movement through the second location within the three-dimensional environment; an electronic device that, in accordance with the first input, moves the first object the second distance away from the first location within the three-dimensional environment;
73. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment a first distance away from the first object within the three-dimensional environment; While displaying the three-dimensional environment, including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, receiving, via one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment by a second distance, the second distance being greater than the first distance; In response to receiving the first input, in accordance with a determination that the first input satisfies a first set of one or more criteria, the first set of criteria including a requirement that the first input correspond to movement through the second location within the three-dimensional environment, moving the first object the first distance away from the first location within the three-dimensional environment in accordance with the first input; in response to a determination that the first input does not satisfy the first set of one or more criteria because the first input does not correspond to movement through the second location within the three-dimensional environment; and moving the first object the second distance away from the first location within the three-dimensional environment in accordance with the first input.
74. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment that is a first distance away from the first object within the three-dimensional environment; means for receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, via one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment by a second distance, the second distance being greater than the first distance; In response to receiving the first input, in accordance with a determination that the first input satisfies a first set of one or more criteria, the first set of criteria including a requirement that the first input correspond to movement through the second location within the three-dimensional environment, moving the first object the first distance away from the first location within the three-dimensional environment in accordance with the first input; in response to a determination that the first input does not satisfy the first set of one or more criteria because the first input does not correspond to movement through the second location within the three-dimensional environment; means for moving the first object the second distance away from the first location in the three-dimensional environment in accordance with the first input.
75. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment that is a first distance away from the first object within the three-dimensional environment; means for receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, via one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment by a second distance, the second distance being greater than the first distance; In response to receiving the first input, in accordance with a determination that the first input satisfies a first set of one or more criteria, the first set of criteria including a requirement that the first input correspond to movement through the second location within the three-dimensional environment, moving the first object the first distance away from the first location within the three-dimensional environment in accordance with the first input; in response to a determination that the first input does not satisfy the first set of one or more criteria because the first input does not correspond to movement through the second location within the three-dimensional environment; means for moving the first object the second distance away from the first location within the three-dimensional environment in accordance with the first input.
76. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 55 to 71.
77. 72. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 55 to 71.
78. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 55 to 71.
79. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising:
72. An information processing device comprising means for carrying out the method of any one of claims 55 to 71.
80. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying, via the display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment; While displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, receiving a first input via the one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, in response to determining that the first input corresponds to moving the first object to a third location within the three-dimensional environment that does not include an object; moving a representation of the first object to the third location within the three-dimensional environment in accordance with the first input; maintaining the display of the first object at the third location after the first input is terminated; in response to a determination that the first input corresponds to a movement of the first object to the second location within the three-dimensional environment and in response to a determination that one or more criteria are satisfied; moving the representation of the first object to the second location within the three-dimensional environment according to the first input; adding the first object to the second object at the second location within the three-dimensional environment.
81. 81. The method of claim 80, wherein, prior to receiving the first input, the first object is contained within a third object at the first location within the three-dimensional environment.
82. Moving the first object away from the first location within the three dimensional environment in accordance with the first input comprises:
82. The method of claim 81, comprising removing the representation of the first object from the third object at the first location within the three dimensional environment in accordance with a first portion of the first input, and moving the representation of the first object within the three dimensional environment in accordance with a second portion of the first input while the third object remains at the first location within the three dimensional environment.
83. 83. The method of claim 81 or 82, further comprising: while moving the representation of the first object away from the first location in the three-dimensional environment in accordance with the first input, displaying, via the display generation component, a second representation of the first object, different from the representation of the first object, within the third object at the first location in the three-dimensional environment.
84. moving the representation of the first object away from the first location within the three-dimensional environment in accordance with the first input, and then in response to detecting an end of the first input, 84. The method of any one of claims 80 to 83, further comprising displaying an animation of the first representation of the first object moving to the first location within the three-dimensional environment in accordance with a determination that the current location of the first object satisfies one or more second criteria, including criteria that are satisfied when the current location within the three-dimensional environment is an invalid location for the first object.
85. in response to detecting an end of the first input after moving the representation of the first object to the third location within the three dimensional environment in accordance with the first input because the third location within the three dimensional environment does not include an object; generating a third object at the third location within the three-dimensional environment; 85. The method of any one of claims 80 to 84, further comprising displaying the first object within the third object at the third location within the three-dimensional environment.
86. The one or more criteria include a criterion that is satisfied when the second object is a valid drop target for the first object, and the method further comprises: after moving the representation of the first object to the second location within the three-dimensional environment in accordance with the first input, such that the first input corresponds to moving the first object to the second location within the three-dimensional environment; upon determining that the one or more criteria are satisfied because the second object is a valid drop target for the first object, displaying, via the display generating component, a visual indicator overlaid on the first object indicating that the second object is the valid drop target for the first object; 86. The method of any one of claims 80 to 85, further comprising canceling generation of the third object at the second location within the three-dimensional environment.
87. The one or more criteria include a criterion that is satisfied when the second object is a valid drop target for the first object, and the method further comprises: in response to detecting an end of the first input after moving the representation of the first object to the second location within the three-dimensional environment in accordance with the first input, such that the first input corresponds to moving the first object to the second location within the three-dimensional environment; upon determining that the one or more criteria are not satisfied because the second object is an invalid drop target for the first object; ceasing to display the representation of the first object at the second location within the three-dimensional environment; 87. The method of claim 85 or 86, further comprising canceling generation of the third object at the second location within the three-dimensional environment.
88. 88. The method of claim 87, further comprising: after moving the representation of the first object to the second location within the three-dimensional environment in accordance with the first input, where the first input corresponds to moving the first object to the second location within the three-dimensional environment, in accordance with the determination that the one or more criteria are not satisfied because the second object is an invalid drop target for the first object, displaying, via the display generation component, a visual indicator overlaid on the first object indicating that the second object is an invalid drop target for the first object.
89. 89. The method of any one of claims 80 to 88, wherein the second object includes a three-dimensional drop zone for receiving the object when the second object is a valid drop target for the object, the drop zone extending from the second object toward a viewpoint of the user within the three-dimensional environment.
90. Before the first object reaches the drop zone of the second object in accordance with the first input, the first object has a first size within the three-dimensional environment, and the method further comprises:
90. The method of claim 89, further comprising, in response to moving the representation of the first object into the drop zone of the second object as part of the first input, resizing the first object in the three dimensional environment to have a second size different from the first size.
91. the three-dimensional environment includes a fifth object at a fourth location within the three-dimensional environment, the fifth object including a sixth object, and the method further comprising: receiving, while displaying the three-dimensional environment containing the fifth object, the fifth object including the sixth object at the fourth location within the three-dimensional environment, via the one or more input devices, a second input corresponding to a request to move the fifth object to the second location within the three-dimensional environment; In response to receiving the second input, In response to a determination that the fifth object has a distinct characteristic, adding the sixth object to the second object at the second location within the three-dimensional environment; 91. The method of any one of claims 80 to 90, further comprising ceasing to display the fifth object in the three-dimensional environment.
92. In response to receiving the second input, In response to a determination that the fifth object does not have the individual characteristic, 92. The method of claim 91, further comprising adding the fifth object, which includes the sixth object contained in the fifth object, to the second object at the second location within the three-dimensional environment.
93. the three-dimensional environment includes a fifth object at a fourth location within the three-dimensional environment, the fifth object including a sixth object, and the method further comprising: while displaying the three-dimensional environment including the fifth object that encompasses the sixth object at the fourth location within the three-dimensional environment; displaying, via the display generation component, one or more interface elements associated with the fifth object at the fourth location within the three-dimensional environment in accordance with a determination that one or more second criteria are satisfied, the second criteria including criteria that are satisfied when a gaze of a user of the electronic device is directed toward the fifth object; 93. The method of any one of claims 80 to 92, further comprising: ceasing to display the one or more interface elements associated with the fifth object in accordance with a determination that the one or more second criteria are not satisfied.
94. 94. The method of claim 93, wherein the one or more second criteria include a criterion that is met when a predetermined portion of the user of the electronic device has a distinct pose and that is not met when the predetermined portion of the user of the electronic device does not have the distinct pose.
95. In response to receiving the first input, In accordance with the determination that the first input corresponds to moving the first object to the third location within the three-dimensional environment that does not include the object, the first object is displayed at the third location along with a first distinct user interface element associated with the first object for moving the first object within the three-dimensional environment; 95. The method of any one of claims 80 to 94, wherein, in accordance with the determination that the first input corresponds to moving the first object to the second location within the three-dimensional environment and in accordance with a determination that the one or more criteria are satisfied, the first object is displayed at the second location without the first separate user interface element associated with the first object for moving the first object within the three-dimensional environment.
96. receiving, via the one or more input devices, a second input corresponding to a request to move the first object within the three-dimensional environment while displaying the first object at the third location using the first individual user interface element associated with the first individual object to move the first individual object within the three-dimensional environment; While receiving the second input, ceasing to display the first discrete user interface element; 96. The method of claim 95, further comprising: moving a representation of the first object within the three-dimensional environment according to the second input.
97. The three-dimensional environment includes a third object at a fourth location within the three-dimensional environment, and the method further comprises: while displaying the three-dimensional environment including the second object that encompasses the first object at the second location within the three-dimensional environment and the third object at the fourth location within the three-dimensional environment, receiving, via the one or more input devices, a second input including a first portion of a second input corresponding to a request to move the first object away from the second object at the second location within the three-dimensional environment, the first portion being followed by a second portion of the second input; while receiving the first portion of the second input, moving the representation of the first object away from the second object at the second location within the three dimensional environment in accordance with the first portion of the second input; In response to detecting an end of the second portion of the second input, maintain a display of the first object within the second object at the second location within the three-dimensional environment in accordance with a determination that the second portion of the second input corresponds to moving the first object to the fourth location within the three-dimensional environment and a determination that one or more second criteria are not satisfied because the third object is not a valid drop target for the first object; 97. The method of any one of claims 80 to 96, further comprising: maintaining a display of the first object within the second object at the second location within the three-dimensional environment in accordance with a determination that the second portion of the second input corresponds to movement of the first object to the second location within the three-dimensional environment.
98. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment; While displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, receive, via one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, in response to determining that the first input corresponds to moving the first object to a third location within the three-dimensional environment that does not include an object; moving a representation of the first object to the third location within the three-dimensional environment in accordance with the first input; maintaining the display of the first object at the third location after the first input is terminated; in response to a determination that the first input corresponds to a movement of the first object to the second location within the three-dimensional environment and in response to a determination that one or more criteria are satisfied; moving the representation of the first object to the second location within the three-dimensional environment according to the first input; An electronic device that adds the first object to the second object at the second location within the three-dimensional environment.
99. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment; While displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, receiving, via one or more input devices, a first input corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, in response to determining that the first input corresponds to moving the first object to a third location within the three-dimensional environment that does not include an object; moving a representation of the first object to the third location within the three-dimensional environment in accordance with the first input; maintaining the display of the first object at the third location after the first input is terminated; in response to a determination that the first input corresponds to a movement of the first object to the second location within the three-dimensional environment and in response to a determination that one or more criteria are satisfied; moving the representation of the first object to the second location within the three-dimensional environment according to the first input; and adding the first object to the second object at the second location within the three-dimensional environment.
100. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment; means for receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, in response to determining that the first input corresponds to moving the first object to a third location within the three-dimensional environment that does not include an object; moving a representation of the first object to the third location within the three-dimensional environment in accordance with the first input; maintaining the display of the first object at the third location after the first input is terminated; in response to a determination that the first input corresponds to a movement of the first object to the second location within the three-dimensional environment and in response to a determination that one or more criteria are satisfied; moving the representation of the first object to the second location within the three-dimensional environment according to the first input; means for adding the first object to the second object at the second location within the three-dimensional environment.
101. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a three-dimensional environment including a first object at a first location within the three-dimensional environment and a second object at a second location within the three-dimensional environment; means for receiving, while displaying the three-dimensional environment including the first object at the first location within the three-dimensional environment and the second object at the second location within the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the first object away from the first location within the three-dimensional environment; In response to receiving the first input, in response to determining that the first input corresponds to moving the first object to a third location within the three-dimensional environment that does not include an object; moving a representation of the first object to the third location within the three-dimensional environment in accordance with the first input; maintaining the display of the first object at the third location after the first input is terminated; in response to a determination that the first input corresponds to a movement of the first object to the second location within the three-dimensional environment and in response to a determination that one or more criteria are satisfied; moving the representation of the first object to the second location within the three-dimensional environment according to the first input; means for adding the first object to the second object at the second location within the three-dimensional environment.
102. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 80 to 97.
103. 98. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 80 to 97.
104. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 80 to 97.
105. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising:
98. An information processing apparatus comprising means for carrying out the method of any one of claims 80 to 97.
106. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying, via the display generation component, a three-dimensional environment including a plurality of objects, the three-dimensional environment including a first object and a second object different from the first object; detecting, while displaying the three-dimensional environment, via the one or more input devices, a first input corresponding to a request to move the plurality of objects to a first location within the three-dimensional environment, followed by termination of the first input; While detecting the first input, moving representations of the plurality of objects in the three-dimensional environment together to the first location in accordance with the first input; and separately placing the first object and the second object within the three-dimensional environment in response to detecting the end of the first input.
107. 107. The method of claim 106, wherein the first input comprises a first movement of a discrete portion of a user of the electronic device, the first movement corresponding to the movement to the first location within the three-dimensional environment, followed by the termination of the first input.
108. detecting the termination of the first input and separately placing the first object and the second object within the three-dimensional environment, followed by detecting a second input via the one or more input devices corresponding to a request to move the first object to a second location within the three-dimensional environment; In response to receiving the second input, moving the first object to the second location within the three-dimensional environment without moving the second object within the three-dimensional environment; detecting the termination of the first input and separately placing the first object and the second object within the three-dimensional environment, followed by detecting a third input via the one or more input devices corresponding to a request to move the second object to a third location within the three-dimensional environment; 108. The method of claim 106 or 107, further comprising: in response to receiving the third input, moving the second object to the second location within the three-dimensional environment without moving the first object within the three-dimensional environment.
109. 109. The method of any one of claims 106 to 108, further comprising, while detecting the first input, displaying, within the three dimensional environment, a visual indication of the number of objects included in the plurality of objects toward which the first input is directed.
110. 110. The method of claim 109, wherein, during detection of the first input, the plurality of objects are arranged in a discrete arrangement having positions within the discrete arrangement associated with an order, and the visual indication of the number of objects included in the plurality of objects is displayed at a discrete location for a discrete object within the plurality of objects located at a primary position within the discrete arrangement.
111. detecting a second input via the one or more input devices corresponding to a request to add a third object to the plurality of objects while displaying the visual indication of the number of objects included in the plurality of objects at the respective location for the respective object located at the primary position within the respective arrangement; In response to detecting the second input, adding the third object to the individual arrangement, the third object being placed in the primary position within the individual arrangement rather than the individual object; 111. The method of claim 110, further comprising: displaying the visual indication of the number of objects included in the plurality of objects at the distinct location relative to the third object.
112. the visual indication of the number of objects included in the plurality of objects is displayed in a location based on the individual objects within the plurality of objects; pursuant to a determination that the distinct object is a two-dimensional object, the visual indication is displayed on the two-dimensional object; 112. A method according to any one of claims 109 to 111, wherein, following a determination that the distinct object is a three-dimensional object, the visual indication is displayed on a boundary of a bounding volume containing the distinct object.
113. detecting, while displaying the plurality of objects along with the visual indication of the number of the objects included in the plurality of objects, a second input via the one or more input devices corresponding to a request to move the plurality of objects to a third object within the three-dimensional environment; While detecting the second input, moving the representations of the plurality of objects to the third object; 113. The method of any one of claims 109 to 112, further comprising: updating the visual indication to indicate a number of objects included in the plurality of objects for which the third object is a valid drop target, wherein the number of objects included in the plurality of objects is different from the number of objects included in the plurality of objects for which the third object is a valid drop target.
114. 114. The method of any one of claims 106 to 113, further comprising: in response to detecting the end of the first input, determining that the first location where the end of the first input was detected is an empty space within the three-dimensional environment, and positioning the plurality of objects differently based on the first location such that the plurality of objects are positioned at different distances from a viewpoint of the user.
115. 115. The method of claim 114, wherein separately arranging the plurality of objects comprises arranging the plurality of objects in a spiral pattern within the three-dimensional environment.
116. 116. The method of claim 115, wherein the radius of the spiral pattern increases as a function of the distance of the user from the viewpoint.
117. 117. The method of any one of claims 114 to 116, wherein the plurality of separately positioned objects are confined to a volume defined by the first location within the three-dimensional environment.
118. While detecting the first input, the plurality of objects are arranged in a separate arrangement having positions within the separate arrangement associated with an order, and separately arranging the plurality of objects based on the first location includes: placing a discrete object in a primary position within the discrete arrangement at the first location; and placing other objects in the plurality of objects at different locations in the three-dimensional environment based on the first location.
119. In response to detecting the end of the first input, in response to determining that the first location where the termination of the first input was detected is an empty space within the three-dimensional environment; displaying the first object within the three-dimensional environment using a first user interface element for moving the first object within the three-dimensional environment; displaying the second object within the three-dimensional environment using a second user interface element for moving the second object within the three-dimensional environment; 119. The method of any one of claims 106 to 118, further comprising: displaying the first object and the second object in the three-dimensional environment without displaying the first user interface element and the second user interface element in accordance with a determination that the first location where the end of the first input was detected includes a third object.
120. in response to detecting the end of the first input and in accordance with a determination that the first location at which the end of the first input was detected includes a third object; pursuant to a determination that the third object is an invalid drop target for the first object, displaying, via the display generation component, an animation of the representation of the first object moving to a location within the three-dimensional environment where the first object was located when the first input was detected; 120. The method of any one of claims 106 to 119, further comprising: in accordance with a determination that the third object is an invalid drop target for the second object, displaying, via the display generation component, an animation of the representation of the second object moving to a location in the three-dimensional environment where the second object was located when the first input was detected.
121. after detecting the termination of the first input and after separately placing the first object and the second object within the three-dimensional environment, detecting a second input via the one or more input devices corresponding to a request to select one or more of the plurality of objects for movement within the three-dimensional environment; In response to detecting the second input, selecting the plurality of objects for movement within the three-dimensional environment in accordance with a determination that the second input was detected within a respective time threshold of detecting the termination of the first input; 121. The method of any one of claims 106 to 120, further comprising: discontinuing selection of the plurality of objects for movement within the three-dimensional environment in accordance with a determination that the second input is detected after the respective time threshold of detecting the termination of the first input.
122. in accordance with a determination that the plurality of objects were moving at a velocity greater than a velocity threshold when the end of the first input was detected, the individual time threshold is a first time threshold; 122. The method of claim 121, wherein the individual time threshold is a second time threshold that is less than the first time threshold in accordance with a determination that the plurality of objects were moving at a velocity less than the velocity threshold when the end of the first input was detected.
123. detecting a second input via the one or more input devices while detecting the first input and moving the representations of the plurality of objects together in accordance with the first input, the second input including detecting a separate portion of a user of the electronic device performing a separate gesture while a gaze of the user is directed toward a third object in the three-dimensional environment, the third object not being included in the plurality of objects; 123. The method of any one of claims 106 to 122, further comprising: in response to detecting the second input, adding the third object to the plurality of objects being moved together in the three-dimensional environment in accordance with the first input.
124. detecting a second input via the one or more input devices while detecting the first input and moving representations of the plurality of objects together in accordance with the first input, comprising detecting a discrete portion of the user of the electronic device performing a discrete gesture while a user's gaze is directed toward a third object within the three-dimensional environment, the third object not being included in the plurality of objects, followed by movement of the discrete portion of the user corresponding to movement of the third object to a current location of the plurality of objects within the three-dimensional environment; 124. The method of any one of claims 106 to 123, further comprising: in response to detecting the second input, adding the third object to the plurality of objects being moved together in the three-dimensional environment in accordance with the first input.
125. detecting the first input and, while moving the representations of the plurality of objects together in accordance with the first input, detecting a second input via the one or more input devices corresponding to a request to add a third object to the plurality of objects; In response to detecting the second input, in accordance with a determination that the third object is a two-dimensional object, adding the third object to the plurality of objects being moved together in the three-dimensional environment in accordance with the first input; 125. The method of any one of claims 106 to 124, further comprising adjusting at least one dimension of the third object based on a corresponding dimension of the first object in the plurality of objects.
126. the first object is a two-dimensional object; the second object is a three-dimensional object; prior to detecting the first input, the first object has a smaller size in the three-dimensional environment than the second object; 126. The method of any one of claims 106 to 125, wherein the representation of the second object has a smaller size than the representation of the first object while the multiple objects move together in accordance with the first input.
127. While detecting the first input, the plurality of objects are arranged in an individual arrangement having positions within the individual arrangement associated with an order; the first object is a three-dimensional object and the second object is a two-dimensional object; 127. A method according to any one of claims 106 to 126, wherein the first object is displayed in a preferred position relative to the second object in the individual arrangement, regardless of whether the first object was added to the plurality of objects before the second object was added to the plurality of objects or whether the first object was added to the plurality of objects after the second object was added to the plurality of objects.
128. detecting a second input via the one or more input devices while detecting the first input, the second input including detecting a separate portion of the user of the electronic device performing a separate gesture while the user's gaze is directed at the plurality of objects; 128. The method of any one of claims 106 to 127, further comprising, in response to detecting the second input, removing the individual object from the plurality of objects such that the individual object is no longer moved within the three-dimensional environment in accordance with the first input.
129. the plurality of objects includes a third object; the first object and the second object are two-dimensional objects; the third object is a three-dimensional object; While detecting the first input, a representation of the first object is displayed parallel to a representation of the second object; 129. A method according to any one of claims 106 to 128, wherein a predetermined surface of the representation of the third object is displayed perpendicular to the representations of the first and second objects.
130. While detecting the first input, the first object is displayed at a first distance from a viewpoint of a user of the electronic device and at a first relative orientation with respect to a viewpoint of a user of the electronic device; 130. The method of any one of claims 106 to 129, wherein the second object is displayed at a second distance from the user's viewpoint and at a second relative orientation different from the first relative orientation to the user's viewpoint.
131. 131. The method of any one of claims 106 to 130, wherein, while detecting the first input, the plurality of objects act as drop targets for one or more other objects in the three-dimensional environment.
132. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions displaying, via a display generation component, a three-dimensional environment including a plurality of objects, including a first object and a second object different from the first object; detecting, while displaying the three-dimensional environment, via one or more input devices, a first input corresponding to a request to move the plurality of objects to a first location within the three-dimensional environment, followed by termination of the first input; while detecting the first input, moving representations of the plurality of objects in the three-dimensional environment together to the first location in accordance with the first input; In response to detecting the termination of the first input, the electronic device separately positions the first object and the second object within the three-dimensional environment.
133. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: displaying, via a display generation component, a three-dimensional environment including a plurality of objects, the three-dimensional environment including a first object and a second object different from the first object; detecting, while displaying the three-dimensional environment, via one or more input devices, a first input corresponding to a request to move the plurality of objects to a first location within the three-dimensional environment, followed by termination of the first input; While detecting the first input, moving representations of the plurality of objects in the three-dimensional environment together to the first location in accordance with the first input; and separately placing the first object and the second object within the three-dimensional environment in response to detecting the end of the first input.
134. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a three-dimensional environment including a plurality of objects, the plurality of objects including a first object and a second object different from the first object; means for detecting, while displaying the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the plurality of objects to a first location within the three-dimensional environment, followed by termination of the first input; means for moving representations of the plurality of objects in the three-dimensional environment together to the first location in accordance with the first input while detecting the first input; means for separately placing the first object and the second object within the three-dimensional environment in response to detecting the termination of the first input.
135. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a three-dimensional environment including a plurality of objects, the plurality of objects including a first object and a second object different from the first object; means for detecting, while displaying the three-dimensional environment, a first input via one or more input devices corresponding to a request to move the plurality of objects to a first location within the three-dimensional environment, followed by termination of the first input; means for moving representations of the plurality of objects in the three-dimensional environment together to the first location in accordance with the first input while detecting the first input; means for separately placing the first object and the second object within the three-dimensional environment in response to detecting the end of the first input.
136. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 106 to 131.
137. 132. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 106 to 131.
138. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 106 to 131.
139. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising:
132. An information processing apparatus comprising means for carrying out the method of any one of claims 106 to 131.
140. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying, via the display generation component, a three-dimensional environment including a first object and a second object; While displaying the three-dimensional environment, detecting a first input directed toward the first object via the one or more input devices, the first input comprising a request to throw the first object into the three-dimensional environment at a discrete velocity and a discrete direction; In response to detecting the first input, moving the first object to the second object within the three-dimensional environment in accordance with a determination that the first input satisfies one or more criteria, including criteria that were satisfied when a second object was currently targeted by the user when the request to throw the first object was detected; in accordance with a determination that the first input does not satisfy the one or more criteria because the second object was not currently targeted by the user when the request to throw the first object was detected, moving the first object to a distinct location within the three dimensional environment other than the second object, the distinct location being on a path within the three dimensional environment determined based on the distinct velocity and the distinct direction of the request to throw the first object within the three dimensional environment.
141. 141. The method of claim 140, wherein the first input comprises a movement of a discrete portion of the user of the electronic device corresponding to the discrete velocity and the discrete direction.
142. 142. The method of claim 140 or 141, wherein the second object is targeted based on the user of the electronic device's gaze being directed toward the second object during the first input.
143. 143. The method of any one of claims 140 to 142, wherein moving the first object towards the second object comprises moving the first object towards the second object at a rate based on the respective rates of the first inputs.
144. moving the first object to the second object includes moving the first object within the three-dimensional environment based on a first physics model; 144. The method of any one of claims 140 to 143, wherein moving the first object to the distinct location comprises moving the first object within the three-dimensional environment based on a second physics model that is different from the first physics model.
145. 145. The method of claim 144, wherein moving the first object based on the first physics model includes limiting movement of the first object to a first maximum velocity established by the first physics model, and wherein moving the first object based on the second physics model includes limiting movement of the first object to a second maximum velocity established by the second physics model, the second maximum velocity being different from the first maximum velocity.
146. 146. The method of claim 145, wherein the first maximum speed is greater than the second maximum speed.
147. 147. The method of any one of claims 144 to 146, wherein moving the first object based on the first physics model includes limiting movement of the first object to a first minimum velocity established by the first physics model, and moving the first object based on the second physics model includes limiting movement of the first object to a second minimum velocity established by the second physics model, the second minimum velocity being different from the first minimum velocity.
148. the first minimum velocity is greater than a minimum velocity requirement for the first input to be identified as a throwing input; 148. The method of clause 147, wherein the second minimum velocity corresponds to the minimum velocity requirement for the first input to be identified as the throwing input.
149. 149. The method of any one of claims 140 to 148, wherein the second object is targeted based on the user's gaze of the electronic device being directed toward the second object during the first input and the individual direction of the first input being directed toward the second object.
150. moving the first object to the second object within the three-dimensional environment includes displaying a first animation of the first object moving through the three-dimensional environment to the second object; 150. The method of any one of claims 140 to 149, wherein moving the first object to the distinct location within the three-dimensional environment includes displaying a second animation of the first object moving through the three-dimensional environment to the distinct location.
151. The first animation of the first object moving through space in the three-dimensional environment to the second object comprises: a first portion in which the first animation of the first object corresponds to movement along the path within the three-dimensional environment determined based on the individual speed and the individual direction of the request to throw the first object; and a second portion following the first portion, during which the first animation of the first object corresponds to movement along a different path towards the second object.
152. 152. The method of claim 150 or 151, wherein the second animation of the first object moving through space within the three-dimensional environment to the individual location includes an animation of the first object corresponding to movement along the path to the individual location within the three-dimensional environment determined based on the individual speed and the individual direction of the request to throw the first object.
153. a minimum velocity requirement for the first input to be identified as a throwing input pursuant to a determination that the second object was not currently being targeted by the user when the request to throw the first object was detected, the minimum velocity requirement being a first velocity requirement; 153. The method of any one of claims 140 to 152, wherein, pursuant to a determination that the second object was currently targeted by the user when the request to throw the first object was detected, the minimum speed requirement for the first input to be identified as the throwing input is a second speed requirement that is different from the first speed requirement.
154. 154. The method of any one of claims 140 to 153, wherein moving the first object to the second object within the three-dimensional environment comprises moving the first object to a location within the second object determined based on a line of sight of the user of the electronic device.
155. In response to a determination that the second object includes a content placement area that includes a plurality of different valid locations of the first object and the user's line of sight is directed toward the content placement area within the second object, according to a determination that the user's line of sight is directed toward a first valid location of the plurality of different valid locations relative to the first object, the location within the second object determined based on the user's line of sight is the first valid location; 155. The method of claim 154, wherein, in accordance with a determination that the user's line of sight is directed toward a second valid location among the plurality of different valid locations for the first object that is different from the first valid location, the location within the second object determined based on the user's line of sight is the second valid location.
156. Moving the first object to the second object within the three-dimensional environment comprises:
156. The method of any one of claims 140 to 155, comprising, in accordance with a determination that the second object includes an input field that includes a valid location of the first object, moving the first object to the input field.
157. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions displaying, via a display generation component, a three-dimensional environment including the first object and the second object; While displaying the three-dimensional environment, detecting a first input directed at the first object via one or more input devices, the first input comprising a request to throw the first object into the three-dimensional environment at a discrete velocity and a discrete direction; In response to detecting the first input, moving the first object to the second object within the three-dimensional environment in accordance with a determination that the first input satisfies one or more criteria, including criteria that were satisfied when a second object was currently targeted by the user when the request to throw the first object was detected; and, in accordance with a determination that the first input does not satisfy the one or more criteria because the second object was not currently targeted by the user when the request to throw the first object was detected, moves the first object to a distinct location within the three-dimensional environment other than the second object, the distinct location being on a path within the three-dimensional environment determined based on the distinct velocity and the distinct direction of the request to throw the first object within the three-dimensional environment.
158. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: displaying, via a display generation component, a three-dimensional environment including the first object and the second object; While displaying the three-dimensional environment, detecting a first input directed toward the first object via one or more input devices, the first input comprising a request to throw the first object into the three-dimensional environment at a discrete velocity and a discrete direction; In response to detecting the first input, moving the first object to the second object within the three-dimensional environment in accordance with a determination that the first input satisfies one or more criteria, including criteria that were satisfied when a second object was currently targeted by the user when the request to throw the first object was detected; in accordance with a determination that the first input does not satisfy the one or more criteria because the second object was not currently targeted by the user when the request to throw the first object was detected, moving the first object to a distinct location within the three dimensional environment other than the second object, the distinct location being on a path within the three dimensional environment determined based on the distinct velocity and the distinct direction of the request to throw the first object within the three dimensional environment.
159. 1. An electronic device comprising: one or more processors; Memory and means for displaying, via a display generation component, a three-dimensional environment including the first object and the second object; means for detecting, while displaying the three-dimensional environment, a first input directed at the first object via one or more input devices, the first input comprising a request to throw the first object into the three-dimensional environment at a discrete velocity and a discrete direction; In response to detecting the first input, moving the first object to the second object within the three-dimensional environment in accordance with a determination that the first input satisfies one or more criteria, including criteria that were satisfied when a second object was currently targeted by the user when the request to throw the first object was detected; and means for, in accordance with a determination that the first input does not satisfy the one or more criteria because the second object was not currently targeted by the user when the request to throw the first object was detected, moving the first object to a distinct location within the three-dimensional environment other than the second object, the distinct location being on a path within the three-dimensional environment determined based on the distinct velocity and the distinct direction of the request to throw the first object within the three-dimensional environment.
160. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising: means for displaying, via a display generation component, a three-dimensional environment including the first object and the second object; means for detecting, while displaying the three-dimensional environment, a first input directed at the first object via one or more input devices, the first input comprising a request to throw the first object into the three-dimensional environment at a discrete velocity and a discrete direction; In response to detecting the first input, moving the first object to the second object within the three-dimensional environment in accordance with a determination that the first input satisfies one or more criteria, including criteria that were satisfied when a second object was currently targeted by the user when the request to throw the first object was detected; and means for, in accordance with a determination that the first input does not satisfy the one or more criteria because the second object was not currently targeted by the user when the request to throw the first object was detected, moving the first object to a distinct location within the three-dimensional environment other than the second object, the distinct location being on a path within the three-dimensional environment determined based on the distinct velocity and the distinct direction of the request to throw the first object within the three-dimensional environment.
161. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 140 to 156.
162. 157. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 140 to 156.
163. 1. An electronic device comprising: one or more processors; Memory and and means for performing the method of any one of claims 140 to 156.
164. 1. An information processing apparatus for use in an electronic device, the information processing apparatus comprising:
157. An information processing apparatus comprising means for carrying out the method of any one of claims 140 to 156.
165. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 18.