Devices, methods, and graphical user interfaces for interacting with a three-dimensional environment
Improved computer systems with display generation components and input devices, utilizing touch-sensitive displays and gesture recognition, address inefficiencies in virtual and augmented reality interactions, enhancing user experience and reducing energy consumption.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-08-27
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods and interfaces for interacting with virtual and augmented reality environments are cumbersome, inefficient, and complex, leading to increased cognitive burden and energy consumption, particularly in battery-powered devices.
The implementation of improved computer systems with display generation components and input devices, including touch-sensitive displays, cameras, eye-tracking, and hand-tracking, to enhance user interaction through reduced and intuitive input methods, such as gesture recognition and visual feedback, and dynamic adjustment of virtual and physical object interactions.
This approach reduces the number and complexity of user inputs, enhances user safety and satisfaction, and improves the efficiency of human-machine interaction in three-dimensional environments by making interactions more intuitive and less energy-intensive.
Smart Images

Figure 0007849423000001 
Figure 0007849423000002 
Figure 0007849423000003
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 62 / 907,614, filed on September 28, 2019, and U.S. Patent Application No. 17 / 030,219, filed on September 23, 2020, and is a continuation of U.S. Patent Application No. 17 / 030,219, filed on September 23, 2020.
Technical Field
[0002] The present disclosure generally relates to a computer system having a display generation component, including but not limited to, an electronic device that provides virtual and mixed reality experiences via a display, and one or more input devices that provide a computer-generated experience.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch-sensing surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include digital images, videos, text, icons, and virtual objects such as buttons and other graphics.
[0004] However, the methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems where manipulating virtual objects is complex and error-prone impair the user's cognitive burden and detract from the experience in virtual / augmented reality environments. In addition, these methods are unnecessarily time-consuming, thereby wasting energy. The latter problem is particularly critical in battery-powered devices. [Overview of the Initiative]
[0005] Therefore, there is a need for computer systems with improved methods and interfaces to provide users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or types of user input by helping the user understand the connection between the inputs provided and the device response to those inputs, thereby generating a more efficient human-machine interface.
[0006] The above-mentioned defects and other problems relating to the user interface for a computer system having a display generation component and one or more input devices are mitigated or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generation component, the output devices include one or more tactile output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI or the user's body as captured by a camera and other motion sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally include non-temporary computer-readable storage media or other computer program products configured to be executed by one or more processors.
[0007] There is a need for electronic devices with improved methods and interfaces for interacting with three-dimensional environments. Such methods and interfaces can complement or replace conventional methods for interacting with three-dimensional environments. Such methods and interfaces reduce the number, extent, and / or types of user input, resulting in a more efficient human-machine interface.
[0008] There is a need for electronic devices with improved methods and interfaces for generating computer-generated environments. Such methods and interfaces can complement or replace conventional methods for generating computer-generated environments. Such methods and interfaces can generate more efficient human-machine interfaces, allowing users to have greater control over the device, and enabling users to use devices that are safer, require less cognitive effort, and have an improved user experience.
[0009] In some embodiments, the method is performed in a computer system including a display generation component and one or more input devices, and the method is The process includes: displaying a virtual object at a first spatial position in a three-dimensional environment; detecting a first hand movement performed by the user while the virtual object is displayed at the first spatial position in the three-dimensional environment; in response to the detection of the first hand movement performed by the user, performing a first action in accordance with the first hand movement without moving the virtual object away from the first spatial position, according to a determination that the first hand movement satisfies a first gesture criterion; displaying a first visual indication that the virtual object has entered reconstruction mode, according to a determination that the first hand movement satisfies a second gesture criterion; detecting a second hand movement performed by the user while the virtual object is displayed with the first visual indication that the virtual object has entered reconstruction mode; and in response to the detection of the second hand movement performed by the user, moving the virtual object from the first spatial position to a second spatial position in accordance with the second hand movement, according to a determination that the second hand movement satisfies a first gesture criterion.
[0010] According to some embodiments, the method is performed in a computer system including a display generation component and one or more input devices, and displays a three-dimensional scene via the display generation component, in which a first virtual object is displayed with a first value of a first display characteristic corresponding to a first part of the first virtual object and a second value of a first display characteristic different from a second part of the first virtual object, and the second value of the first display characteristic is different from the first value of the first display characteristic, and includes at least a first virtual object at a first position and a first physical plane at a second position separate from the first position, and while displaying the three-dimensional scene including the first virtual object and the first physical plane, the display generation The process includes generating a first visual effect at a second location in a three-dimensional scene via a component, wherein generating the first visual effect modifies the visual appearance of a first portion of a first physical plane in the three-dimensional scene according to a first value of a first display characteristic corresponding to a first portion of a first virtual object, and modifies the visual appearance of a second portion of a first physical plane in the three-dimensional scene according to a second value of a first display characteristic corresponding to a second portion of a first virtual object, wherein the visual appearance of the first portion of the first virtual object and the visual appearance of the second portion of the first physical plane are modified differently depending on the difference between the first and second values of the first display characteristics of the first and second portions of the first physical plane.
[0011] According to some embodiments, the method is performed in a computer system including a display generation component and one or more input devices, and comprises displaying a three-dimensional scene via the display generation component, wherein the three-dimensional scene includes a first set of physical elements and a first amount of virtual elements, and the first set of physical elements includes at least physical elements corresponding to a first class of physical objects and physical elements corresponding to physical elements corresponding to a second class of physical objects; detecting a sequence of two or more user inputs while displaying the three-dimensional scene including the first amount of virtual elements via the display generation component; and in response to detecting a sequence of two or more user inputs, continuously increasing the amount of virtual elements displayed in the three-dimensional scene according to the sequence of two or more user inputs, and in response to detecting a first user input in a sequence of two or more user inputs, determining that the first user input satisfies a first criterion. The method includes displaying a three-dimensional scene with at least a first subset of one or more physical elements from a first set and a second quantity of virtual elements, wherein the second quantity of virtual elements occupies a larger portion of the three-dimensional scene than the first quantity of virtual elements, including a first portion of the three-dimensional scene occupied by a first class of physical elements before the detection of the first user input; and, in response to the detection of a second user input in a sequence of two or more user inputs, the method includes displaying a three-dimensional scene with at least a second subset of one or more physical elements from a first set and a third quantity of virtual elements, wherein the third quantity of virtual elements occupies a larger portion of the three-dimensional scene than the second quantity of virtual elements, including a first portion of the three-dimensional scene occupied by a first class of physical elements before the detection of the first user input and a second portion of the three-dimensional scene occupied by a second class of physical elements before the detection of the second user input, according to the determination that the second user input follows the first user input and satisfies a first criterion.
[0012] According to some embodiments, the method is performed in a computer system including a display generation component and one or more input devices, and the method includes displaying a three-dimensional scene via the display generation component, which includes at least a first physical object having at least a first physical surface, and the position of each of the first physical object or its representation in the three-dimensional scene corresponds to the position of each of the first physical object in the physical environment surrounding the display generation component; detecting that a first interaction criterion has been met while the three-dimensional scene is being displayed, which includes a first criterion that is met when a first level of user interaction between the user and the first physical object is detected; and in response to the detection that the first interaction criterion has been met, The method includes: displaying a first user interface at a location corresponding to the location of a first physical face of a first physical object in a three-dimensional scene via a demonstration generation component; detecting that a second interaction criterion is met while the first user interface is displayed at a location corresponding to the location of a first physical face or its representation of the first physical object in a three-dimensional scene, the second interaction criterion being met when a second level of user interaction higher than a first level of user interaction between the user and the first physical object is detected; and, in response to the detection that the second interaction criterion has been met, replacing the display of the first user interface with a display of the second user interface at the location corresponding to the location of a first physical face or its representation of the first physical object in a three-dimensional scene.
[0013] According to some embodiments, the method is performed on a computer system including a display generation component and one or more input devices, and includes: displaying a three-dimensional scene via the display generation component, wherein the three-dimensional scene includes at least a first physical object having a first physical surface and at least a first virtual object having a first virtual surface; detecting a request to activate a voice-based virtual assistant while displaying the three-dimensional scene including the first physical object and the first virtual object; activating a voice-based virtual assistant configured to receive voice commands in response to the detection of the request to activate the voice-based virtual assistant; displaying a visual representation of the voice-based virtual assistant in the three-dimensional scene, including displaying the visual representation of the voice-based virtual assistant with a first set of values of first display characteristics of the visual representation; and modifying the visual appearance of at least a portion of the first physical surface of the first physical object and at least a portion of the first virtual surface of the first virtual object according to a first set of values of first display characteristics of the visual representation of the voice-based virtual assistant.
[0014] According to some embodiments, the computer system includes a display generation component (e.g., a display, projector, head-mounted display, etc.), one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, one or more processors, and memory for storing one or more programs, the one or more programs being configured to be executed by the one or more processors, and the one or more programs including instructions for performing or causing to perform any of the operations described herein. According to some embodiments, a non-temporary computer-readable storage medium having instructions stored therein, the instructions causing the device to perform or perform any of the operations described herein when executed by a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), and optionally one or more tactile output generators. According to some embodiments, a graphical user interface of a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, one or more sensors that optionally detect the intensity of contact with the touch-sensitive surface), one or more tactile output generators that optionally detect one or more tactile output generators, memory, and one or more processors that execute one or more programs stored in memory, includes one or more elements that are displayed in any of the methods described herein, and these elements are updated in response to input as described in any of the methods described herein. According to some embodiments, the computer system includes a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, one or more sensors that optionally detect the intensity of contact with the touch-sensitive surface), one or more tactile output generators that optionally detect one or more tactile output generators, and means for performing or causing to perform any of the operations described herein.According to some embodiments, an information processing device for use in a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensing surface, optionally one or more sensors for detecting the intensity of contact with the touch-sensing surface), and optionally one or more tactile output generators, includes means for performing or causing to perform any of the operations described herein.
[0015] Therefore, computer systems having display generation components are provided with improved methods and interfaces that enhance the effectiveness, efficiency, and user safety and satisfaction of such computer systems by interacting with a three-dimensional environment and simplifying the user's use of the computer system when interacting with a three-dimensional environment. Such methods and interfaces can complement or replace conventional methods for interacting with a three-dimensional environment and facilitating the user's use of the computer system when interacting with a three-dimensional environment.
[0016] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The functions and advantages described herein are not exhaustive, and many additional functions and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention. [Brief explanation of the drawing]
[0017] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.
[0018] [Figure 1]A block diagram showing the operating environment of a computer system for providing a CGR experience according to some embodiments.
[0019] [Figure 2] A block diagram showing a controller of a computer system configured to manage and adjust a user's CGR experience according to some embodiments.
[0020] [Figure 3] A block diagram showing a display generation component of a computer system configured to provide a visual component of a CGR experience to a user according to some embodiments.
[0021] [Figure 4] A block diagram showing a hand tracking unit of a computer system configured to capture a user's gesture input according to some embodiments.
[0022] [Figure 5] A block diagram showing an eye tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.
[0023] [Figure 6] A flowchart showing a grint-assisted gaze tracking pipeline according to some embodiments.
[0024] <০০০০০৯৪>A block diagram showing user interaction with a computer-generated three-dimensional environment (e.g., including reconstruction and other interactions) according to some embodiments. [Figure 7B] A block diagram showing user interaction with a computer-generated three-dimensional environment (e.g., including reconstruction and other interactions) according to some embodiments.
[0025] [Figure 7C] This block diagram shows several embodiments of methods for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between physical and virtual objects). [Figure 7D] This block diagram shows several embodiments of methods for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between physical and virtual objects). [Figure 7E] This block diagram shows several embodiments of methods for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between physical and virtual objects). [Figure 7F] This block diagram shows several embodiments of methods for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between physical and virtual objects).
[0026] [Figure 7G] This block diagram shows a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion of the computer-generated experience based on user input), according to several embodiments. [Figure 7H] This block diagram shows a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion of the computer-generated experience based on user input), according to several embodiments. [Figure 7I] This block diagram shows a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion of the computer-generated experience based on user input), according to several embodiments. [Figure 7J]This block diagram shows a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion of the computer-generated experience based on user input), according to several embodiments. [Figure 7K] This block diagram shows a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion of the computer-generated experience based on user input), according to several embodiments. [Figure 7L] This block diagram shows a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion of the computer-generated experience based on user input), according to several embodiments.
[0027] [Figure 7M] This block diagram shows a method, according to several embodiments, for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment). [Figure 7N] This block diagram shows a method, according to several embodiments, for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment). [Figure 7O] This block diagram shows a method, according to several embodiments, for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment). [Figure 7P] This block diagram shows a method, according to several embodiments, for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment). [Figure 7Q]This block diagram shows a method, according to several embodiments, for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment). [Figure 7R] This block diagram shows a method, according to several embodiments, for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment).
[0028] [Figure 7S] This block diagram shows a method for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. [Figure 7T] This block diagram shows a method for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. [Figure 7U] This block diagram shows a method for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. [Figure 7V] This block diagram shows a method for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. [Figure 7W] This block diagram shows a method for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. [Figure 7X]This block diagram shows a method for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments.
[0029] [Figure 8] This is a flowchart of methods for interacting with a computer-generated three-dimensional environment (including, for example, reconstruction and other interactions) according to several embodiments.
[0030] [Figure 9] This is a flowchart of a method for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between physical and virtual objects) according to several embodiments.
[0031] [Figure 10] This is a flowchart of a method for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of the computer-generated experience based on user input), according to several embodiments.
[0032] [Figure 11] This is a flowchart illustrating a method for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment) according to several embodiments.
[0033] [Figure 12] This is a flowchart of a method for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. [Modes for carrying out the invention]
[0034] This disclosure relates to user interfaces that provide a computer-generated reality (CGR) experience to a user, in several embodiments.
[0035] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.
[0036] In some embodiments, a computer system allows a user to interact with virtual objects in a computer-generated three-dimensional environment by using various gesture inputs. A first predetermined gesture (e.g., a swipe gesture, tap gesture, pinch gesture, and drag gesture) causes the computer system to perform a first action corresponding to the virtual object, while the same predetermined gesture moves the virtual object in the computer-generated three-dimensional environment from one position to another in the computer system when combined with a special modification gesture (e.g., a reconfiguration gesture) (e.g., immediately after, simultaneously with, or after the reconfiguration gesture). Specifically, in some embodiments, a predetermined reconfiguration gesture causes the virtual object to enter a reconfiguration mode. While in reconfiguration mode, the object can be moved from one position to another in the computer-generated environment in response to a first gesture configured to trigger a first type of interaction with the virtual object when the virtual object is not in reconfiguration mode (e.g., activating the virtual object, navigating within the virtual object, or rotating the virtual object). In some embodiments, the reconfiguration gesture is not part of a gesture that moves the virtual object, and the virtual object optionally remains in reconfiguration mode after entering it in response to detection of a previous reconfiguration mode. While the virtual object is in reconfiguration mode, the computer system optionally responds to other gesture inputs directed to the computer-generated environment without causing the virtual object to exit reconfiguration mode. The computer system moves the virtual object according to a first set of gestures, each configured to trigger a first type of interaction with the virtual object when the virtual object is not in reconfiguration mode. Visual indications of virtual objects that have entered and remain in reconfiguration mode are provided to help the user understand the computer-generated environment and the internal status of the virtual object and to provide appropriate inputs to achieve the desired result.By using a special reconfiguration gesture to put virtual objects into reconfiguration mode, utilizing gestures that typically reconfigure the environment and trigger other actions that move virtual objects, and providing visual indications of virtual objects that enter and remain in reconfiguration mode in response to a special reconfiguration gesture, the number, range, and / or nature of user input is reduced, resulting in a more efficient human-machine interface.
[0037] In some embodiments, the computer system generates a three-dimensional environment that includes both physical objects (e.g., appearing in the three-dimensional environment through transparent or translucent portions of display-generated components or in a camera view of the physical environment) and virtual objects (e.g., user interface objects, computer-generated virtual objects that mimic physical objects, and / or objects that do not have physical counterparts in the real world). The computer system generates mimicked visual interactions between virtual and physical objects according to mimicked physical light propagation principles. Specifically, light emitted from virtual objects (e.g., including luminance, color, hue, temporal variation, spatial patterns, etc.) appears to illuminate both physical and virtual objects in the environment. The computer system generates mimicked illumination and shadows on various parts of physical surfaces and various parts of virtual surfaces, which are brought about by virtual light emitted from virtual objects. The illumination and shadows are generated considering the physical light propagation principles, as well as the spatial position of the virtual objects relative to other physical and virtual surfaces in the environment, the mimicked physical properties of the virtual surfaces (e.g., surface texture, optical properties, shape, and dimensions, etc.), and the actual physical properties of the physical surfaces (e.g., surface texture, optical properties, shape, and dimensions, etc.). Light emitted from different parts of a virtual object affects different parts of other virtual objects and different parts of other physical objects in the environment due to differences in their position and physical properties. By generating realistic and detailed visual interactions between virtual and physical objects, and by making virtual and physical objects react similarly to lighting from virtual objects, computer systems can make three-dimensional environments more realistic, helping users become more familiar with computer-generated three-dimensional environments and reducing user errors when interacting with them.
[0038] In some embodiments, the user provides the computer system with a sequence of two or more predetermined inputs to continuously increase the level of immersion in the computer-generated experience provided by the computer system. When the user positions the computer system's display generation component in a given location relative to the user (e.g., placing a display in front of their eyes or a head-mounted device on their head), the user's view of the real world is obscured by the display generation component, and the content presented by the display generation component dominates the user's view. At times, the user benefits from a more gradual and controlled process of transitioning from the real world to the computer-generated experience. Thus, when displaying content to the user through the display generation component, the computer system displays a pass-through portion that includes representations of at least a portion of the real world surrounding the user, and gradually increases the amount of virtual elements that replace the physical elements visible through the display generation component. Specifically, in response to each sequential input in a sequence of two or more user inputs, different classes of physical elements are removed from the view and replaced by newly displayed virtual elements (e.g., extensions of existing or newly added virtual elements). The gradual transition to and from an immersive environment controlled by user input is intuitive and natural for the user, improving the user experience and comfort when using a computer system for computer-generated immersive experiences. Dividing the physical elements that are replaced as a whole in response to each input into different classes of physical elements further reduces the total number of user inputs required to transition to an immersive computer-generated environment, while allowing user control over multiple gradual transitions.
[0039] In some embodiments, when a computer system displays a three-dimensional environment including physical objects (for example, the physical objects are visible through a display-generating component) (for example, visible through a transparent pass-through portion of the display-generating component as a camera view of the physical environment indicated by the display-generating component, or as a virtual representation of the physical objects in a simulated reality environment rendered by the display-generating component). The physical objects have physical surfaces (for example, planar or smooth surfaces). When the level of interaction between the physical objects and the user is at a first predetermined level, the computer system displays a first user interface at a position corresponding to the location of the physical objects in the three-dimensional environment (for example, so that the first user interface appears to overlap or stand on the physical surface). When the level of interaction between the physical objects and the user is at a second level, for example, a higher level than the first level of interaction, the computer system displays a second user interface that replaces the first user interface at a position corresponding to the location of the physical objects in the three-dimensional environment (for example, so that the second user interface appears to overlap or stand on the physical surface). The second user interface provides more information and / or functionality associated with physical objects compared to the first user interface. The computer system allows the user to interact with the first and second user interfaces using various means for receiving information and controlling the first physical object. This technology enables the user to interact with physical objects with the help of more information and control provided at locations within the computer-generated environment. The locations of interaction within the computer-generated environment correspond to the physical locations of physical objects in the real world.By adjusting the information and level of control (e.g., provided to different user interfaces) according to the detected level of interaction between the user and the physical object, the computer system reduces user confusion and errors when the user interacts with the computer-generated environment without providing unnecessary information or cluttering the computer-generated three-dimensional environment. This technique also allows, in some embodiments, the user to utilize nearby physical surfaces to remotely control physical objects. In some embodiments, the user can control physical objects remotely or control information about physical objects to make the user's interaction with the physical object and / or the three-dimensional environment more efficient.
[0040] In some embodiments, the computer system generates a three-dimensional environment that includes both physical objects (e.g., appearing in the three-dimensional environment through transparent or translucent portions of display-generated components or in a camera view of the physical environment) and virtual objects (e.g., user interface objects, computer-generated virtual objects that mimic physical objects, and / or objects that do not have physical counterparts in the real world). The computer system also provides a voice-based virtual assistant. When the voice-based virtual assistant is activated, the computer system displays a visual representation of the activated virtual assistant. The computer system also modifies the appearance of the physical and virtual objects in the environment, as well as the background of the user's field of view or the peripheral area of the screen, according to the display characteristics values of the visual representation of the virtual assistant. Specifically, the light emitted from the visual representation of the virtual assistant (e.g., including luminance, color, hue, temporal variation, spatial pattern, etc.) appears to illuminate both the physical and virtual objects in the environment, and optionally, the background of the user's or peripheral area of the screen's field of view. The computer system generates mimicking illumination and shadows on various parts of the physical surface and various parts of the virtual surface brought about by the virtual light emitted from the visual representation of the virtual assistant. Lighting and shadows are generated considering the physical principles of light propagation, as well as the spatial position of the virtual assistant's visual representation relative to other physical and virtual surfaces within the computer-generated environment, the simulated physical properties of the virtual surfaces (e.g., surface texture, optical properties, shape, and dimensions), and the actual physical properties of the physical surfaces (e.g., surface texture, optical properties, shape, and dimensions). The lighting effects associated with the virtual assistant provide the user with continuous and dynamic feedback regarding the state of the voice-based virtual assistant (e.g., active or paused, listening, and / or responding).Computer systems can make computer-generated three-dimensional environments more realistic and informative by generating realistic and detailed visual interactions between the visual representation of a virtual assistant and other virtual and physical objects within the computer-generated environment. This helps users become more familiar with the computer-generated three-dimensional environment and reduces user errors when interacting with it.
[0041] Figures 1-6 illustrate exemplary computer systems for providing a CGR experience to a user. Figures 7A-7B are block diagrams illustrating user interaction with a computer-generated three-dimensional environment (including, for example, reconstruction and other interactions) according to several embodiments. Figures 7C-7F are block diagrams illustrating methods for generating a computer-generated three-dimensional environment (including, for example, mimicking visual interactions between physical and virtual objects) according to several embodiments. Figures 7G-7L are block diagrams illustrating methods for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion in the computer-generated experience based on user input) according to several embodiments. Figures 7M-7R are block diagrams illustrating methods for facilitating user interaction with a computer-generated environment (including, for example, controlling a device or interacting with a computer-generated environment by utilizing interaction with a physical surface) according to several embodiments. Figures 7S-7X are block diagrams illustrating methods for generating a computer-generated three-dimensional environment (including, for example, mimicking visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. Figure 8 is a flowchart of several embodiments of methods for interacting with a computer-generated three-dimensional environment (including, for example, reconstruction and other interactions). Figure 9 is a flowchart of several embodiments of methods for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between physical and virtual objects). Figure 10 is a flowchart of several embodiments of methods for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of the computer-generated experience based on user input). Figure 11 is a flowchart of several embodiments of methods for facilitating user interaction with a computer-generated environment (for example, controlling a device using interaction with a physical surface or interacting with a computer-generated environment).Figure 12 is a flowchart of a method for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. The user interfaces in Figures 7A to 7X are used to illustrate the processes in Figures 8 to 12.
[0042] In some embodiments, as shown in Figure 1, the CGR experience is provided to the user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, within a head-mounted device or handheld device).
[0043] When describing a CGR experience, various terms are used to refer individually to several related but distinct environments that the user perceives and / or interacts with (for example, using inputs detected by the computer system 101, which causes the computer system generating the CGR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the CGR experience). The following is a subset of these terms.
[0044] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.
[0045] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In CGR, a subset of a person's bodily movements or their representations are tracked, and in response, one or more properties of one or more virtual objects simulated within the CGR environment are modulated to behave according to at least one law of physics. For example, a CGR system may detect a person's head rotation and, in response, modulate the graphic content and sound field presented to the person in a similar manner to how such views and sounds would change in a physical environment. Depending on the circumstances (e.g., for reasons of accessibility), the modulation of the properties(s) of the virtual object(s) in the CGR environment may be done in response to a representation of bodily movement (e.g., a voice command). A person may perceive and / or interact with CGR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may perceive and / or interact with an audio object that creates a 3D or spatially expansive audio environment, providing the perception of a point source in 3D space. In another example, an audio object may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, a person may perceive and / or interact with only audio objects.
[0046] Examples of CGR include virtual reality and mixed reality.
[0047] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment contains multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment by simulating the presence of a person within the computer-generated environment and / or by simulating a subset of a person's bodily movements within the computer-generated environment.
[0048] Mixed Reality: A mixed reality (MR) environment, in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input, refers to a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects). On a virtual continuum, a mixed reality environment is any place between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track the position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.
[0049] Examples of mixed reality include augmented reality and augmented virtual reality.
[0050] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through images or videos of the physical environment. As used herein, videos of the physical environment shown on an opaque display are referred to as “pass-through videos,” and it means that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, thereby allowing a person to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to an imitation environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the image sensor. As another example, the representation of the physical environment may be transformed by graphically altering (e.g., enlarging) a portion of it, thereby making the altered portion a modified version that represents the original captured image but is not photorealistic. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.
[0051] Augmented Virtual: An Augmented Virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, but people with faces might be realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.
[0052] Hardware: There are many different types of electronic systems that enable people to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to the human eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination thereof. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the human retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces. In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience.In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., physical setup / environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is coupled to a display generation component 120 (e.g., an HMD, display, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within a housing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0053] In some embodiments, the display generation component 120 is configured to provide the user with a CGR experience (e.g., at least the visual component of the CGR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0054] According to some embodiments, the display generation component 120 provides the user with a CGR experience while the user is virtually and / or physically present in scene 105.
[0055] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., the head or hand). Thus, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with CGR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the CGR content response is displayed via the HMD. Similarly, a user interface showing interaction with CGR content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the interaction is triggered by the movement of the HMD relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).
[0056] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features for the sake of simplification are not shown so as not to obscure more suitable embodiments of the exemplary embodiments disclosed herein.
[0057] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0058] In some embodiments, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0059] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and CGR experience module 240.
[0060] The operating system 230 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). For this purpose, in various embodiments, the CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.
[0061] In some embodiments, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 in Figure 1, and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 242 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0062] In some embodiments, the tracking unit 244 is configured to map scene 105, at least display generation component 120 relative to scene 105 in Figure 1, and optionally track the position of one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, the processing unit 244 includes a hand tracking unit 243 and / or an eye tracking unit 245. In some embodiments, the hand tracking unit 243 is configured to track the position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to scene 105 in Figure 1, relative to the display generation component 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 243 is described in more detail below with reference to Figure 4. In some embodiments, the eye-tracking unit 245 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to the CGR content displayed via the display generation component 120. The eye-tracking unit 245 is described in more detail below with reference to Figure 5.
[0063] In some embodiments, the adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0064] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral devices 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0065] While the data acquisition unit 242, tracking unit 244 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., a controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, tracking unit 244 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.
[0066] Furthermore, Figure 2 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary from embodiment to embodiment and in some embodiments will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0067] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting embodiments, the HMD120 may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional inward and / or outward image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0068] In some embodiments, one or more communication buses 304 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0069] In some embodiments, one or more CGR displays 312 are configured to provide the user with a CGR experience. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, the HMD 120 includes a single CGR display. In another embodiment, the HMD 120 includes a CGR display for each of the user's eyes. In some embodiments, one or more CGR displays 312 can present MR or VR content. In some embodiments, one or more CGR displays 312 can present MR or VR content.
[0070] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.
[0071] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and CGR presentation module 340.
[0072] The operating system 330 includes procedures for handling various basic system services and procedures for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to the user via one or more CGR displays 312. Thus, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.
[0073] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0074] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. For this purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0075] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (for example, a 3D map of a mixed reality scene or a map of a physical environment on which computer-generated objects can be placed to generate computer-generated reality) based on media content data. For this purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0076] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0077] Although the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0078] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary from embodiment to embodiment and in some embodiments will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0079] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by the hand tracking unit 243 (Figure 2) to track the position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 in Figure 1 (e.g., relative to a part of the physical environment surrounding the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined for the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0080] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.
[0081] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which in turn drives the display generation component 120. For example, a user can interact with the software running on the controller 110 by moving their hand 408 and changing the hand's orientation.
[0082] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the hand tracking device 440 may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.
[0083] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D positions of the user's wrist and fingertips.
[0084] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The posture estimation function described herein may be interleaved with the motion tracking function, so that patch-based posture estimation is performed only once every two (or more) frames, while tracking is used to detect changes in posture that occur over the remaining frames. Posture, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 in response to the posture and / or gesture information, or perform other functions.
[0085] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 440, but some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the hand tracking device 402, or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (for example, in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.
[0086] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values to identify and segment image components (i.e., adjacent pixel groups) that have features of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.
[0087] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the position and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.
[0088] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 245 (Figure 2) to track the position and movement of the user's gaze relative to the scene 105 or to the CGR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating CGR content for user viewing and a component for tracking the user's gaze relative to the CGR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or CGR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is optionally used in combination with a head-mounted display generation component, rather than being a head-mounted device. In some embodiments, the eye-tracking device 130 is optionally part of a non-head-mounted display generation component, rather than being a head-mounted device.
[0089] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0090] As shown in Figure 5, in some embodiments, the eye-tracking device 130 includes at least one eye-tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visual light to pass through. The eye-tracking device 130 optionally captures images of the user's eye (e.g., as a video stream captured at 60 to 120 frames per second (fps)), analyzes the images to generate eye-tracking information, and communicates the eye-tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a corresponding eye-tracking camera and light source.
[0091] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil position, central visual position, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a glint-assisted method to determine the user's current visual axis and viewpoint relative to the display.
[0092] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece (one or more) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes (one or more) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display or projector of a handheld device) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the user's eye(s) 592 (e.g., as shown at the bottom of Figure 5).
[0093] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses gaze tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's gaze on the display 510 based on the gaze tracking input 542 obtained from the eye-tracking camera 540. The gaze estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0094] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the CGR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can utilize the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.
[0095] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) mounted on a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circular pattern around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0096] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the position and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wide field of view (FOV) and a camera 540 with a narrow FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0097] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality (including, for example, virtual reality and / or mixed reality) applications to provide users with computer-generated reality (including, for example, virtual reality, augmented reality and / or augmented virtual reality) experiences.
[0098] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in a tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in a tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.
[0099] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which is initiated at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0100] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.
[0101] At 640, if the process proceeds from element 410, the current frame is analyzed and the pupil and glint are tracked, based in part on preceding information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to confirm that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.
[0102] Figure 6 is intended to serve as an example of an eye-tracking technique that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking techniques that currently exist or may be developed in the future may be used in the computer system 101 to provide the user with a CGR experience in various embodiments, either in place of or in combination with the glint-assisted eye-tracking technique described herein.
[0103] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, it should be understood that each example is compatible with the input device or method described in the other example and can be used optionally. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, it should be understood that each example is compatible with the output device or method described in the other example and can be used optionally. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with the method described in the other example and can be used optionally. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each embodiment. User interface and related processes
[0104] Here, we focus on embodiments of the user interface ("UI") and related processes that may be performed in a computer system such as a portable multifunction device or head-mounted device, which may include a display generation component, one or more input devices, and (optionally) one or a camera.
[0105] Figures 7A and 7B are block diagrams illustrating user interactions with a computer-generated three-dimensional environment (including, for example, reconstruction and other interactions) according to several embodiments. Figures 7A and 7B are used to illustrate processes described later, including the process shown in Figure 8.
[0106] In some embodiments, the input gestures described with reference to Figures 7A-7B are detected by analyzing data and signals captured by a sensor system (e.g., sensor 190 in Figure 1, image sensor 314 in Figure 3). In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras such as a motion RGB camera, an infrared camera, a depth camera, etc.). For example, one or more imaging sensors are components of a computer system (e.g., computer system 101 in Figure 1 (e.g., portable electronic device 7100 or HMD as shown in Figures 7A-7B)) that includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4 (e.g., a touchscreen display that functions as a display and a touch-sensing surface, a stereoscopic display, a display with a pass-through portion, etc.)) or provide data to the computer system. In some embodiments, one or more imaging sensors include one or more rear cameras on the side of the device opposite to the device's display. In some embodiments, the input gestures are detected by a sensor system of a head-mounted system (e.g., a VR headset that includes a stereoscopic display that provides a left image for the user's left eye and a right image for the user's right eye). For example, one or more cameras, which are components of a head-mounted system, are mounted on the front and / or bottom of the head-mounted system. In some embodiments, one or more imaging sensors are positioned in the space in which the head-mounted system is used (e.g., arranged around the head-mounted system at various positions in a room) so that the imaging sensors capture images of the head-mounted system and / or the user of the head-mounted system. In some embodiments, input gestures are detected by a sensor system of a head-up device (e.g., a head-up display, a car windshield capable of displaying graphics, a window capable of displaying graphics, or a lens capable of displaying graphics). For example, one or more imaging sensors are mounted on the interior of a car.In some embodiments, the sensor system includes one or more depth sensors (e.g., a sensor array). For example, the one or more depth sensors include one or more light-based (e.g., infrared) sensors and / or one or more acoustic-based (e.g., ultrasonic) sensors. In some embodiments, the sensor system includes one or more signal emitters, such as light emitters (e.g., infrared emitters) and / or sound emitters (e.g., ultrasonic emitters). For example, while light (e.g., light from an infrared light emitter array having a predetermined pattern) is projected onto a hand (e.g., a hand 7200 as described with reference to Figures 7A-7B), an image of the hand under light illumination is captured by one or more cameras, and the captured image is analyzed to determine the position and / or configuration of the hand. In contrast to using signals from touch-sensitive surfaces or other direct contact or proximity-based mechanisms, determining input gestures using signals from an image sensor directed at the hand allows users to freely choose whether to perform large movements or remain relatively still when providing input gestures with their hands, without experiencing the constraints imposed by specific input devices or input areas.
[0107] In some embodiments, a plurality of user interface objects 7208, 7210, and 7212 (e.g., within a menu or dock, or independently of each other) are displayed in a computer-generated three-dimensional environment (e.g., a virtual environment or a mixed reality environment). The plurality of user interface objects are optionally displayed floating above physical objects in space or in the three-dimensional environment. Each user interface object optionally has one or more corresponding actions that can be performed in the three-dimensional environment or that cause an action in the physical environment communicating with the computer system (e.g., controlling another device (e.g., a speaker or smart lamp) communicating with device 7100). In some embodiments, the user interface objects 7208, 7210, and 7212 are displayed by the computer system's display (e.g., device 7100 (Figures 7A-7B) or HMD) together with (e.g., overlaid or replaced) at least a portion of a view of the physical environment captured by one or more rear cameras of the computer system (e.g., device 7100). In some embodiments, user interface objects 7208, 7210, and 7212 are displayed on a transparent or translucent display of a computer system (e.g., a head-up display, or HMD) through which the physical environment is visible. In some embodiments, user interface objects 7208, 7210, and 7212 are displayed on a user interface that includes a pass-through portion surrounded by virtual content (e.g., a transparent or translucent portion through which the physical surroundings are visible, or a portion that displays a camera view of the surrounding physical environment). In some embodiments, user interface objects 7208, 7210, and 7212 are displayed in a virtual reality environment (e.g., floating in virtual space or overlapping a virtual surface).
[0108] In some embodiments, the representation of hand 7200 is visible in the virtual reality environment (e.g., an image of hand 7200 captured by one or more cameras is rendered in the virtual reality setting). In some embodiments, a representation of hand 7200' (e.g., a cartoon version of hand 7200) is rendered in the virtual reality setting. In some embodiments, hand 7200 or its representation is invisible in the virtual reality environment (e.g., omitted). In some embodiments, device 7100 (Figure 7C) is invisible in the virtual reality environment (e.g., when device 7100 is an HMD). In some embodiments, an image of device 7100 or a representation of device 7100 is visible in the virtual reality environment.
[0109] In some embodiments, one or more of the user interface objects 7208, 7210, and 7212 are application launch icons (e.g., actions to perform actions to launch the corresponding application and actions to display the corresponding quick action menu for each application). In some embodiments, one or more of the user interface objects 7208, 7210, and 7212 are controls to perform various actions within the application (e.g., increasing volume, decreasing volume, playing, pausing, fast forwarding, rewinding, initiating communication with a remote device, ending communication with a remote device, transmitting communication with a remote device, starting a game, etc.). In some embodiments, one or more of the user interface objects 7208, 7210, and 7212 are each representation (e.g., avatars) of a user of a remote device (e.g., to perform actions to initiate communication with each user of the remote device). In some embodiments, one or more of the user interface objects 7208, 7210, and 7212 are representations (e.g., thumbnails, two-dimensional images, or album covers) of media items (e.g., images, virtual objects, audio files, and / or video files). For example, by activating a user interface object that is a representation of an image, the image is displayed (e.g., at a location corresponding to a surface detected by one or more cameras and displayed in a computer-generated reality view (e.g., at a location corresponding to a surface in a physical environment, or at a location corresponding to a surface displayed in a virtual space)). By navigating within a user interface object that is an album (e.g., a music album, a picture album, a flipbook album, etc.), the currently playing or displayed item is switched to another item in the album.
[0110] As shown in Figure 7A, two distinct actions are performed on user interface objects 7208, 7210, and 7212 in the three-dimensional environment in response to different types of gesture inputs provided by hand 7200, while the reconfiguration mode is not activated for any of the user interface objects.
[0111] In Figures 7A(a-1) and 7A(a-3), the thumb of hand 7200 performs a tap gesture by moving along the vertical axis to contact the side of the index finger and then moving upward away from the side of the index finger. The tap gesture is performed while the current selection indicator (e.g., a selector object, or a movable visual effect such as highlighting an object by changing its outline or appearance) is located on the user interface object 7208 and indicates the currently selected status of the user interface object 7208. In some embodiments, in response to detecting a tap input by hand 7200, the computer system (e.g., device 7100) performs a first action to display a virtual object 7202 (e.g., activate the user interface object 7208) (e.g., as part of the user interface of an application represented by the user interface object 7208, or as content represented by the user interface object 7208). The visual appearance of the user interface object 7208 indicates that the first action has been performed (e.g., it is activated but not moving).
[0112] In Figure 7A(a-1), and then in Figures 7A(a-4) to 7A(a-5), hand 7200 performs a drag gesture by moving the thumb laterally after it touches the side of the index finger. The drag gesture is performed while the current selection indicator (e.g., a selector object, or a movable visual effect such as highlighting an object by changing its outline or appearance) is located on user interface object 7208 and indicates that user interface object 7208 is currently selected. In some embodiments, in response to detecting a drag input by hand 7200, a computer system (e.g., device 7100) performs a second action on user interface object 7208 (e.g., navigating away from user interface object 7208 towards user interface object 7210, or navigating within user interface object 7208). The visual appearance of the user interface object indicates that the second action is being performed (e.g., navigation has occurred within or away from the content of the user interface object, but the object has not moved in the three-dimensional environment).
[0113] Figure 7B shows a scenario that is entirely different from the scenario shown in Figure 7A, in that a reconstruction gesture is performed (for example, in combination with other gesture inputs (e.g., the gesture shown in Figure 7A)), and as a result, the three-dimensional environment is reconstructed (for example, with the movement of the user interface object 7208 in the three-dimensional environment).
[0114] As shown in the sequence in Figures 7B(a-1) to 7B(a-4), a wrist flick gesture is provided by hand 7200 while user interface object 7208 is currently selected. In this embodiment, the wrist flick gesture is a predetermined reconfiguration gesture that puts the currently selected user interface object into reconfiguration mode. In some embodiments, detecting the wrist flick gesture includes detecting a touchdown of the thumb on the side of the index finger, followed by an upward rotation of the hand around the wrist. Optionally, at the end of the wrist flick gesture, the thumb is lifted from the side of the index finger. While user interface object 7208 is selected, in response to the detection of the wrist flick gesture (e.g., by a previous input or by eye-tracking input focused on user interface object 7208), the computer system (e.g., device 7100) activates the reconfiguration mode of user interface object 7208. The computer system also displays a visual indication to the user that user interface object 7208 is currently in reconfiguration mode. In some embodiments, as shown in Figure 7B(b-3), the user interface object is removed from its original position and optionally displayed with a modified appearance (e.g., becoming semi-transparent, being enlarged, and / or floating) to indicate that the user interface object 7208 is in reconstruction mode. In some embodiments, after the reconstruction gesture is completed, the user interface object 7208 remains in reconstruction mode and the visual indication remains displayed in the three-dimensional environment. In some embodiments, the computer system optionally responds to other user inputs and provides interaction with the three-dimensional environment according to those inputs, while the user interface object 7208 remains in reconstruction mode (e.g., floating in its original position with a modified appearance).In some embodiments, the computer system optionally allows the user to use a second wrist flick gesture to bring another currently selected user interface object (e.g., the user optionally selects another object by gaze or tap input) into reconfiguration mode while user interface object 7208 remains in reconfiguration mode. In some embodiments, the computer system allows the user to look away or navigate to other parts of the three-dimensional environment without moving or interacting with the user interface objects in reconfiguration mode while one or more user interface objects (e.g., user interface object 7208) remain in reconfiguration mode. In some embodiments, in contrast to those shown in Figures 7A(a-4) to 7A(a-5), a subsequent drag gesture (e.g., performed by the thumb touching the side of the index finger and then the hand 7200 moving laterally) can cause the user interface object 7208 in reconfiguration mode to move from its current position to another position in the three-dimensional environment in accordance with the hand movement (e.g., as shown in Figures 7B(a-5) to 7B(a-6)). In some embodiments, moving the user interface object 7208 according to a drag gesture does not cause the user interface object to exit reconfiguration mode. While the user interface object 7208 remains in reconfiguration mode, one or more additional drag gestures may be optionally used to rearrange the user interface object 7208 in the three-dimensional environment. In some embodiments, a predetermined exit gesture (e.g., a downward wrist flick gesture (e.g., a downward wrist flick gesture performed at the end of a drag gesture, or a standalone downward wrist flick gesture that is not part of another gesture)) causes the user interface object 7208 to exit reconfiguration mode.In some embodiments, once the user interface object 7208 exits reconstruction mode, its appearance is restored to its original state and settles into the target position specified by the drag input(s) directed at the user interface object during reconstruction mode.
[0115] As shown in the sequence of Figures 7B(a-5) to 7B(a-6) following Figures 7B(a-1) to 7B(a-2), a wrist flick gesture provided by hand 7200 is the beginning of a compound gesture that ends with a drag gesture provided by hand 7200. The wrist flick gesture is detected while user interface object 7208 is currently selected. In this embodiment, the wrist flick gesture causes the currently selected user interface object to enter reconfiguration mode and move to another location in accordance with the movement of the drag gesture. In some embodiments, after the user interface object (e.g., user interface object 7208) has entered reconfiguration mode, the user interface object may optionally remain in reconfiguration mode after being moved from one location to another in the environment by the drag input.
[0116] In some embodiments, other types of gestures are optionally used as reconfiguration gestures to activate the reconfiguration mode of the currently selected user interface object. In some embodiments, a given gesture is configured to activate the reconfiguration mode of user interface objects of each class in a three-dimensional environment (e.g., bringing multiple user interface objects of the same class (e.g., application icon class, content item class, object class representing physical objects, etc.) into reconfiguration mode), allowing user interface objects of each class to be moved individually or synchronously in the three-dimensional environment in accordance with subsequent movement inputs (e.g., drag inputs). In some embodiments, the computer system activates the reconfiguration mode of a user interface object in response to detecting a tap input (e.g., on a finger or controller) while the user interface object is selected (e.g., by a previous input or gaze input). In some embodiments, the computer system activates the reconfiguration mode of a user interface object in response to detecting a swipe input (e.g., on a finger or controller) while the user interface object is selected (e.g., by a previous input or gaze input).
[0117] In some embodiments, while a user interface object is in reconstruction mode, the computer system displays a visual indicator (e.g., a shadow or semi-transparent image of the user interface object) following the user's gaze or finger movement to indicate the target position of the user interface object in a three-dimensional environment. In response to detecting a subsequent commitment input (e.g., a downward wrist flick gesture or a tap input on a finger or controller), the computer system positions the user interface object at the current position of the visual indicator.
[0118] In some embodiments, the drag input shown in Figures 7A and 7B is replaced by a swipe input on a finger or controller to perform the corresponding function.
[0119] In some embodiments, the movement of user interface objects in a three-dimensional environment mimics the movement of physical objects in the real world and is constrained by virtual and physical surfaces within the three-dimensional environment. For example, when a virtual object is moved in response to a drag input while it is in reconstruction mode, the virtual object slides across physical surfaces represented in the three-dimensional environment and optionally across virtual surfaces within the three-dimensional environment. In some embodiments, the user interface object jumps when switching between physical surfaces represented in the three-dimensional environment.
[0120] In some embodiments, the computer system optionally generates audio output (e.g., continuous or one or more discrete audio outputs) while the user interface object is in reconstruction mode.
[0121] Figures 7C to 7F are block diagrams illustrating methods for generating a computer-generated three-dimensional environment (including, for example, simulating the visual interaction between physical and virtual objects) according to several embodiments. Figures 7C to 7F are used to illustrate the processes described later, including the process shown in Figure 9.
[0122] Figures 7D–7F show exemplary computer-generated environments corresponding to the physical environment shown in Figure 7C. As described herein with reference to Figures 7D–7F, according to some embodiments, the computer-generated environment may optionally be an augmented reality environment including a camera view of the physical environment, or a computer-generated environment displayed on a display such that it is superimposed on a view of the physical environment which is visible through the transparent portion of the display. As shown in Figure 7C, user 7302 is standing in a physical environment (e.g., scene 105) in which a computer system (e.g., computer system 101) is operating (e.g., holding device 7100 or wearing an HMD). In some embodiments, as in the embodiments shown in Figures 7C–7F, device 7100 is a handheld device (e.g., a mobile phone, tablet, or other mobile electronic device) including a display, touch-sensitive display, etc. In some embodiments, device 7100 is a head-up display or head-mounted This represents a wearable headset including a display, which can be optionally replaced. In some embodiments, the physical environment includes one or more physical surfaces and physical objects surrounding the user 7302 (e.g., the walls of a room (e.g., the front wall 7304 and the side walls 7306), the floor 7308, and furniture 7310). In some embodiments, one or more physical surfaces of physical objects in the environment (e.g., the front 8312 of furniture 7310) are visible through the display generation components of the computer system (e.g., on the display of device 7100 or via the HMD).
[0123] In the embodiments shown in Figures 7D-7F, a computer-generated three-dimensional environment corresponding to the physical environment (for example, a portion of the physical environment that is within the field of view of one or more cameras of device 7100 or visible through the transparent portion of the display of device 7100) is displayed on device 7100. The physical environment includes physical objects having representations corresponding to the computer-generated three-dimensional environment shown by the display generation component of the computer system. For example, in a computer-generated environment shown on a display, the front wall 7304 is represented by the front wall representation 7304', the side wall 7306 is represented by the side wall representation 7306', the floor 7308 is represented by the floor representation 7308', the furniture 7310 is represented by the furniture representation 7310', and the front of the furniture 7310 7312 is represented by the front representation 7312' (for example, the computer-generated environment is an augmented reality environment that includes representations 7304', 7306', 7308', 7310', and 7312' of physical objects that are part of the live view of one or more cameras of device 7100, or physical objects that are visible through the transparent portion of the display of device 7100). In some embodiments, the computer-generated environment shown on a display also includes virtual objects. According to some embodiments, as the field of view of the device 7100 with respect to the physical environment changes (for example, as the field of view of the device 7100 or one or more cameras of the device 7100 with respect to the physical environment changes in response to the movement and / or rotation of the device 7100 within the physical environment), the field of view of the computer-generated environment displayed on the device 7100 changes accordingly (including, for example, changes in the field of view of physical surfaces and physical objects (e.g., walls, floors, furniture, etc.)).
[0124] As shown in Figure 7E, a first virtual object (e.g., a virtual window 7332) is displayed at a first position (e.g., a position in the three-dimensional environment corresponding to a position on a side wall 7306 in the physical environment) in response to user input, for example, to add virtual content to the three-dimensional environment. The first virtual object (e.g., a virtual window 7332) has respective spatial relationships with respect to the representations of the physical objects in the three-dimensional environment (e.g., front wall representation 7304', furniture representation 7310', physical surface representation 7312', and floor representation 7308'), which are determined by the respective spatial relationships between the side walls 7306 and other physical objects (e.g., front wall 7304, furniture 7310, physical surface 7312, and floor 7308). As shown in Figure 7E, the first virtual object (e.g., virtual window 7332) is displayed with a first appearance (e.g., having first luminance values and / or color values for the first parts 7332-b and 7332-c of the first virtual object, and second luminance values and / or color values for the second parts 7332-a and 7332-d). In some embodiments, these internal variations of the display characteristics within the various parts of the first virtual object reflect the content displayed in the first virtual object, which may change due to external factors, predetermined conditions, or over time.
[0125] As shown in Figure 7E, the computer system generates a simulated lighting pattern on a representation of a physical object in a three-dimensional environment based on virtual light emitted from various parts of the virtual object 7332. According to some embodiments, the simulated lighting pattern is generated according to the relative spatial position of the virtual object and the representation of the physical object in the three-dimensional environment, as well as the physical properties of the virtual and physical objects (e.g., surface shape, texture, and optical properties). As shown in Figure 7E, the lighting pattern generated on the representation of the physical object observes the simulated physical light propagation principle. For example, the shape, brightness, color, hue, etc., of the lighting pattern (e.g., lighting patterns 7334, 7336, and 7340) on the representation of the physical object (e.g., representations 7304', 7310', 7312', and 7308') simulates the lighting pattern on the physical object (e.g., physical object / surface 7304, 7310, 7312, and 7308) that would have been made by an actual window with similar properties to the virtual window 7332 on the side wall 7306.
[0126] As shown in Figure 7E, in some embodiments, the computer system generates a simulated lighting pattern 7334 for the front wall 7304 by modifying the visual appearance (e.g., luminance and color values) of the first parts 7334-b and 7334-c of the front wall representation 7304' in the three-dimensional scene according to the luminance and color values of the first parts 7332-b and 7332-c of the first virtual object 7332. Similarly, the computer system generates a simulated lighting pattern 7336 for the physical surface 7312 by modifying the visual appearance (e.g., luminance and color values) of the first parts 7336-b and 7336-c of the physical surface representation 7312' in the three-dimensional scene according to the luminance and color values of the first parts 7332-b and 7332-c of the first virtual object 7332. Similarly, the computer system generates a simulated lighting pattern 7340 of the floor 7308 by modifying the visual appearance (e.g., brightness and color values) of the first parts 7340-b and 7340-c of the floor representation 7308' of the three-dimensional scene according to the brightness and color values of the first parts 7332-b and 7332-c of the first virtual object 7332.
[0127] As shown in Figure 7E, the visual appearance of the first part of the physical surface and the visual appearance of the second part of the physical surface are modified differently according to, for example, the simulated spatial relationship between the first virtual object and the various physical surfaces, the real and simulated physical properties of the virtual object and the various physical surfaces, and the differences in luminance and color values in the various parts of the first virtual object.
[0128] As shown in Figure 7E, in addition to adding simulated lighting patterns 7334, 7336, and 7340 to positions in the three-dimensional environment corresponding to the positions of physical surfaces in the physical environment (e.g., the front wall 7304, the physical surface 7312 of the furniture 7310, and the floor 7308), the computer system also generates a simulated shadow 7338 to a position in the three-dimensional environment (e.g., on the floor representation 7338') corresponding to the position of the actual shadow (e.g., on the floor 7308) that would have been cast by the furniture 7310 if it had been illuminated by an actual light source at the same position and characteristics as the virtual object 7332 (e.g., the actual window on the side wall 7306).
[0129] Figure 7F, compared to Figure 7E, shows how dynamic changes in various parts of a virtual object affect the representation of different parts of the physical environment in different ways. For example, the size and internal contents of the first virtual object have changed in Figure 7E from those shown in Figure 7F. Here, the first virtual object is represented as virtual object 7332'. The first parts of the first virtual object 7332-b and 7332-c in Figure 7E correspond to the first parts 7332-b' and 7332-c' in Figure 7F, respectively. The second parts 7332-a and 7332-d in Figure 7E correspond to the second parts 7332-a' and 7332-d' in Figure 7F, respectively. The center positions of the first parts 7332-b' and 7332-c' and the second parts 7332-a' and 7332-d' in Figure 7F have also shifted relative to the center positions shown in Figure 7E. As a result, for many positions on the side wall representation 7306', the luminance and color values of the corresponding positions on the first virtual object 7332 changed (for example, from the values shown in Figure 7E to the values shown in Figure 7F). Similarly, for many positions on the lighting patterns 7334, 7336, and 7340 projected onto representations 7304', 7312', and 7308', the luminance and color values of the lighting patterns also changed (for example, from the values shown in Figure 7E to the values shown in Figure 7F). For example, for a first position on the side wall representation 7306', the luminance and color values of the corresponding positions on the first virtual object (for example, a virtual window or virtual video screen) could be switched from 1 to 0.5 and from yellow to blue, respectively, and for a second position on the side wall representation 7306', the luminance and color values of the corresponding positions on the first virtual object could be switched from 0.5 to 1 and from blue to yellow, respectively. In some embodiments, a change in the size or movement of the first virtual object causes the brightness and color of some positions on the side wall representation 7306' to change because the first virtual object has expanded or moved to those positions, while the brightness and color of some other positions on the side wall representation 7306' to change because the first virtual object has moved or shrunk from those positions.Furthermore, in some embodiments, the direction of light coming from various parts of the first virtual object can also be optionally changed (for example, the light direction changes according to the time of day or according to the scenery shown in the virtual window). As a result, changes in brightness and color at various locations on the first virtual object cause different changes in illumination at various locations on the representation of the nearby physical surface. Various relationships are used to modify the appearance of the representation of the nearby physical surface based on the appearance of the first virtual object.
[0130] As shown in Figure 7F, the first section 7332-b' brings illumination 7334-b' onto the front wall representation 7304' but no illumination onto the front representation 7312', and the illumination 7336-a' provided by the second section 7332-a' covers the area previously covered by the illumination 7334-b provided by the first section 7332-b (Figure 7E). Similarly, the second section 7332-d' brings illumination 7334-d' onto the front wall representation 7304' but no illumination onto the front representation 7312', and the illumination 7336-c' provided by the first section 7332-c' covers the area previously covered by the illumination 7334-d provided by the second section 7332-d (Figure 7E). Similarly, in the front wall representation 7304', some parts that were previously covered by the lighting 7334-c provided by the first part 7332-c and the lighting 7334-a provided by the second part 7332-a are no longer covered by any lighting. Similarly, in the floor representation 7308', because the first virtual object has contracted, some parts that were previously covered by the lighting 7334-c provided by the first part 7332-c and the lighting 7334-a provided by the second part 7332-a are no longer covered by any lighting. Here, some positions on the floor representation 7308' that were previously covered by higher lighting are now covered by lower lighting, and other positions on the floor representation 7308' that were previously covered by lower lighting are now covered by higher lighting. In Figure 7F, the shadow 7338 cast on the floor representation 7308' also appears darker than the shadow 7308 in Figure 7E because the amount of illumination is reduced due to the reduction in the size of the first virtual object 7332.
[0131] In some embodiments, the first virtual object is a virtual window displaying a virtual landscape. The light emitted from the virtual window is based on the virtual landscape displayed in the virtual window. In some embodiments, the virtual window projects a lighting pattern onto a representation of a nearby physical surface in a three-dimensional environment so as to mimic how light from a real window illuminates a nearby physical surface (based on, for example, the spatial relationship between the window and the physical surface, the physical properties of the physical surface, and the physical light propagation principle). In some embodiments, the virtual landscape displayed in the virtual window changes based on parameters such as the time of day, the location of the landscape, and the size of the virtual window.
[0132] In some embodiments, the first virtual object is a virtual screen or hologram displaying a video. As video playback progresses, the virtual light emitted from the virtual screen or hologram changes as the scene in the video changes. In some embodiments, the virtual screen or hologram projects an illumination pattern onto a representation of a nearby physical surface in a three-dimensional environment so as to mimic how light from an actual video screen or hologram illuminates a nearby physical surface (based, for example, on the spatial relationship between the screen or hologram and the physical surface, the physical properties of the physical surface, and the physical light propagation principle).
[0133] In some embodiments, the first virtual object is a virtual assistant, and the light emitted from the virtual assistant changes during different modes of interaction between the user and the virtual assistant. For example, the visual representation of the virtual assistant has a first color and intensity when first activated by the user, changes to a different color when asking or responding to a question, and changes to a different color when performing a task or waiting for the task to be completed or for a response from the user. In some embodiments, the virtual assistant projects a lighting pattern onto representations of nearby physical surfaces in a three-dimensional environment to mimic how light from a real light source illuminates nearby physical surfaces (based on, for example, the spatial relationship between the light source and the physical surface, the physical properties of the physical surface, and the physical light propagation principle). Additional aspects of how the visual representation of the virtual assistant affects the appearance of nearby physical objects and virtual objects in a three-dimensional environment are illustrated with reference to Figures 7S-7X and Figure 12.
[0134] In some embodiments, the computer system also generates virtual reflections and virtual shadows on a representation of a physical surface based on light emitted from virtual objects near the physical surface.
[0135] Figures 7G to 7L are block diagrams illustrating methods, according to several embodiments, for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of immersion in the computer-generated experience based on user input). Figures 7G to 7L are used to illustrate processes described later, including the process shown in Figure 10.
[0136] Figure 7G shows an exemplary computer-generated environment corresponding to a physical environment. As described herein with reference to Figure 7G, the computer-generated environment may be an augmented reality environment or a computer-generated environment displayed on a display, which is superimposed on a view of the physical environment that is visible through the transparent portion of the display. As shown in Figure 7G, user 7302 is present in a physical environment (e.g., scene 105) that operates a computer system (e.g., computer system 101) (e.g., holding device 7100 or wearing an HMD). In some embodiments, as in the embodiment shown in Figure 7G, device 7100 is a handheld device (e.g., a mobile phone, tablet, or other mobile electronic device) including a display, a touch-sensitive display, etc. In some embodiments, device 7100 represents a wearable headset including a head-up display or a head-mounted display, etc., which may be optionally replaced. In some embodiments, the physical environment includes one or more physical surfaces and physical objects surrounding the user (e.g., the walls of a room (represented by front wall representation 7304', side wall representation 7306'), the floor (e.g., represented by floor representation 7308'), furniture (e.g., represented by furniture representation 7310), and the physical surface 7312 of the furniture (e.g., represented by physical surface representation 7312')).
[0137] In the embodiments shown in Figures 7G to 7L, a computer-generated three-dimensional environment corresponding to the physical environment (e.g., a portion of the physical environment that is within the field of view of one or more cameras of device 7100 or visible through the transparent portion of the display of device 7100) is displayed on device 7100. The computer-generated environment displayed on device 7100 is a three-dimensional environment, and according to some embodiments, as the viewpoint of device 7100 with respect to the physical environment changes (e.g., as the field of view of device 7100 or one or more cameras of device 7100 with respect to the physical environment changes in response to the movement and / or rotation of device 7100 within the physical environment), the viewpoint of the computer-generated environment displayed on device 7100 changes accordingly (e.g., including changes in the viewpoints of physical surfaces and physical objects (e.g., walls, floors, furniture, etc.)).
[0138] As shown in Figure 7G, initially, the three-dimensional environment is shown with a first set of physical elements, including representations of a front wall 7304, side walls 7306, floor 7308, and furniture 7310. Optionally, the three-dimensional environment may also include a first quantity of virtual elements. For example, when the three-dimensional environment is first displayed, or when the computer system's display generation component is first turned on or worn on the user's head or in front of the user's eyes, no virtual elements are displayed in the three-dimensional environment, or only a minimal amount of virtual elements are displayed. This allows the user to start with a view of the three-dimensional environment that closely resembles a direct view of the real world, without the display generation component obstructing the user's view.
[0139] As shown in Figures 7G and 7H, the computer system detects a first predetermined gesture input that enhances the immersion of the three-dimensional environment (e.g., a thumb flick or swipe gesture performed by the hand 7200, represented by a representation 7200' on a display generation component, an upward wave gesture in the air, a swipe gesture on a controller, etc.). In response to the detection of the first predetermined gesture, the computer system displays a virtual element 7402 (e.g., a virtual landscape or virtual window) that obstructs the view of the front wall 7304 in the three-dimensional environment (e.g., the virtual element 7402 replaces the display of the representation 7304' of the front wall 7304 on the display, or the virtual element 7402 is displayed in a position that obstructs the view of the front wall 7304 through a previously transparent portion of the display (the portion now displaying the virtual element 7402)). In some embodiments, as shown in Figure 7H, the view of the front wall 7304 is obscured by the display of the virtual element 7402, but the view of the furniture 7310 in front of the front wall 7304 remains unaffected. In other words, a first predetermined gesture causes a first class of physical object or physical surface (e.g., the front wall) to be replaced or obscured by a newly displayed virtual element or a newly displayed portion of an existing virtual element. In some embodiments, an animated transition is displayed to show the virtual element 7402 gradually expanding (e.g., as shown in Figure 7H) or becoming more opaque and saturated to cover or obscure the view of the front wall 7304 (e.g., replacing the representation 7304' in a three-dimensional environment).
[0140] In some embodiments, in response to a first predetermined gesture, the computer system also optionally adds another virtual element (e.g., virtual object 7404) to the three-dimensional environment without replacing the entire class of physical elements. The virtual object 7404 is optionally a user interface object such as a menu (e.g., an application menu, a document, etc.), a control (e.g., a display brightness control, a display focus control, etc.), or another object that can be manipulated by user input or that provides information or feedback to the three-dimensional environment (e.g., a virtual assistant, a document, a media item, etc.). In some embodiments, as shown in Figure 7I, the virtual object 7404 is added to the three-dimensional environment without gaining input focus and / or without being specifically inserted into the three-dimensional environment (e.g., dragged from a menu or drawn by a drawing tool) (e.g., to block off a portion of the floor 7308 or replace a portion of the floor representation 7308'). In some embodiments, the computer system allows the user to individually introduce each virtual element into the three-dimensional environment using the user interface currently provided to the three-dimensional environment (e.g., adding new furniture, throwing virtual confetti into a room), but this type of input does not alter the immersion of the three-dimensional environment and does not replace the view of all classes of physical elements in a single action.
[0141] Figure 7I, following Figure 7H, shows that the front wall 7304 is completely obscured or replaced by the virtual element 7402. The view of the furniture 7310 in front of the front wall 7304 is still shown in the three-dimensional environment. The virtual element 7404 obscures a portion of the floor representation 7308'. The representation 7306' of the side wall 7306 and the representation 7308' of the floor 7308 are visible in the three-dimensional environment after the virtual elements 7402 and 7404 are added to the three-dimensional environment in response to a first predetermined gesture input.
[0142] As shown in Figures 7I and 7J, after detecting a first predetermined gesture input (for example, shown in Figure 7G), the computer system detects a second predetermined gesture input to enhance immersion in the three-dimensional environment (for example, a thumb flick or swipe gesture performed by a hand 7200 represented by a representation 7200' on a display generation component, a swipe gesture on a controller, etc.). In response to detecting the second predetermined gesture, the computer system maintains the display of a virtual element 7402 (for example, a virtual landscape, or virtual window) that obstructs the view of the front wall 7304 in the three-dimensional environment, and displays a virtual element 7406. The virtual element 7406 obstructs the view of the side wall 7306 in the three-dimensional environment (for example, the virtual element 7406 replaces the display of the representation 7306' of the side wall 7306 on the display, or the virtual element 7406 is positioned to obstruct the view of the side wall 7306 through a previously transparent portion of the display (for example, the portion now displaying the virtual element 7406). In Figures 7I to 7J, a second predetermined gesture causes an additional class of physical object or surface (e.g., a side wall) to be replaced or obstructed by a newly displayed virtual element or a newly displayed portion of an existing virtual element. In some embodiments, an animated transition is displayed in which the virtual element 7406 gradually expands or becomes more opaque, covering or obstructing the view of the side wall 7306 (for example, replacing the representation 7306' in the three-dimensional environment).
[0143] Figure 7K, following Figure 7J, shows that the front wall 7304 and side wall 7306 are completely obscured or replaced by virtual elements 7402 and 7406. A view of the furniture 7310 in front of the front wall 7304 is still shown in the three-dimensional environment. Virtual element 7404 obscures a portion of the floor representation 7308'. After virtual elements 7402, 7404, and 7406 are added to the three-dimensional environment in response to the first and second predetermined gesture inputs, the representation 7308' of the floor 7308 is still visible in the three-dimensional environment.
[0144] As shown in Figures 7K and 7L, after detecting first and second predetermined gesture inputs (for example, shown in Figures 7G and 7I), the computer system detects a third predetermined gesture input to enhance immersion in the three-dimensional environment (e.g., a thumb flick or swipe gesture performed by a hand 7200 represented by a representation 7200' on a display generation component, a swipe gesture on a controller, etc.). In response to detecting the third predetermined gesture input, the computer system maintains the display of virtual elements 7402 and 7406 (e.g., virtual landscape or virtual window) that obstruct the view of the front wall 7304 and side wall 7306 in the three-dimensional environment, and displays virtual elements 7408 and 7410. The virtual element 7408 obstructs the view of the floor 7308 in the three-dimensional environment (for example, the virtual element 7408 replaces the display of the representation 7308' of the floor 7308 on the display, or the virtual element 7408 is positioned to obstruct the view of the floor 7306 through a previously transparent portion of the display (for example, the portion now displaying the virtual element 7408). In Figures 7K to 7L, a third predetermined gesture causes an additional class of physical object or surface (e.g., floor) to be replaced or obstructed by a newly displayed virtual element or a newly displayed portion of an existing virtual element. In some embodiments, an animated transition is displayed in which the virtual element 7408 gradually expands or becomes more opaque, covering or obstructing the view of the floor 7308 (for example, replacing the representation 7308' in the three-dimensional environment).
[0145] In some embodiments, in response to a third predetermined gesture, the computer system also optionally adds another virtual element (e.g., virtual element 7410) to the three-dimensional environment without replacing an entire class of physical elements. The virtual element 7410 is optionally a user interface object such as a menu (e.g., an application menu, document, etc.), a control (e.g., a display brightness control, display focus control, etc.), or another object that can be operated by user input or that provides information or feedback to the three-dimensional environment (e.g., a virtual assistant, document, media item, etc.), or a texture that alters the appearance of a physical object (e.g., decorative features, photograph, etc.). In some embodiments, as shown in Figure 7L, the virtual object 7410 is added to the three-dimensional environment (e.g., overlapping a portion of the front 7312 of furniture 7310, or replacing a portion of the physical surface representation 7312').
[0146] In some embodiments, after a series of predetermined gesture-type input gestures to enhance immersion in the three-dimensional environment, an additional number of virtual elements are optionally introduced into the three-dimensional environment to replace or block views of additional classes of physical elements that were previously visible in the three-dimensional environment. In some embodiments, the entire three-dimensional environment is replaced by virtual elements, and the view of the physical world is completely replaced by a view of virtual elements within the three-dimensional environment.
[0147] In some embodiments, virtual elements 7402 and 7406 are virtual windows displayed in place of corresponding portions of the front wall representation 7304' and the side wall representation 7306', respectively. In some embodiments, the light emitted from the virtual windows casts a simulated lighting pattern onto physical surfaces (e.g., floors or furniture) that are still visible or represented in the three-dimensional environment. Further details of the effects of light from virtual elements on surrounding physical surfaces, according to some embodiments, are described with reference to Figures 7C–7F and 9.
[0148] In some embodiments, the content or appearance of virtual elements 7402 and 7406 (e.g., virtual windows or virtual screens) changes in response to additional gesture input (e.g., a horizontal swipe of the hand in the air, or a swipe of the finger in a predetermined direction). In some embodiments, the size of the virtual elements, the position of the virtual scenery displayed within the virtual elements, the media items displayed within the virtual elements, etc., change in response to additional gesture input.
[0149] In some embodiments, the gesture input for increasing or decreasing the immersion of the three-dimensional environment is a vertical swipe gesture in opposite directions (e.g., upward to increase the amount of immersion / virtual elements, and downward to decrease the amount of immersion / virtual elements). In some embodiments, the gesture for changing the content of a virtual element is a horizontal swipe gesture (e.g., a horizontal swipe gesture switches the content displayed on the virtual element backward and / or forward through multiple positions or times).
[0150] In some embodiments, a sequence of first, second, and third predetermined gesture inputs for increasing immersion in a three-dimensional environment can be optionally replaced by a single sequential input to vary the level of immersion. Each sequential portion of the sequential input corresponds to the respective inputs of the first, second, and third predetermined gesture inputs shown in Figures 7G to 7L, according to some embodiments.
[0151] In some embodiments, the floor 7308 or floor representation 7308' remains visible in the three-dimensional environment even when other physical surfaces, such as walls, are replaced or superimposed by virtual elements. This helps ensure that the user feels safe and avoids tripping when navigating the three-dimensional environment by walking around in the physical world.
[0152] In some embodiments, certain pieces of furniture or parts of furniture surfaces remain visible even when other physical surfaces, such as walls and floors, are replaced or superimposed by virtual elements. This helps ensure a natural relationship with the environment when the user is immersed in the three-dimensional environment.
[0153] In this embodiment, in Figures 7G, 7I, and 7K, a representation 7200' of hand 7200 is displayed in a computer-generated environment. The computer-generated environment does not include a representation of the user's right hand (for example, because the right hand is not within the field of view of one or more cameras of device 7100). Furthermore, in some embodiments, for example, in the example shown in Figure 7I where device 7100 is a handheld device, the user can see parts of the surrounding physical environment apart from any representation of the physical environment displayed on device 7100. For example, parts of the user's hand are visible to the user outside the display of device 7100. In some embodiments, device 7100 in these examples can represent and be replaced by a headset having a display (e.g., a head-mounted display) that completely blocks the user's view of the surrounding physical environment. In some such embodiments, no part of the physical environment is directly visible to the user; instead, the physical environment is visible to the user through representations of parts of the physical environment displayed by the device. In some embodiments, the user's hands are invisible to the user directly or via the device's display while the device continuously or periodically monitors the current state of the user's hands to determine whether the user's hands are ready to provide gesture input. In some embodiments, the device displays an indicator indicating whether the user's hands are ready to provide input gestures, provides feedback to the user, and warns the user to adjust the position of their hands if they wish to provide input gestures.
[0154] Figures 7M to 7R are block diagrams illustrating methods for facilitating user interaction with a computer-generated environment (for example, controlling a device by utilizing interaction with a physical surface or interacting with a computer-generated environment) according to several embodiments. Figures 7M to 7R are used to illustrate processes described later, including the process shown in Figure 11.
[0155] Figure 7N shows an exemplary computer-generated environment corresponding to the physical environment shown in Figure 7M. As described herein with reference to Figures 7M-7R, according to some embodiments, the computer-generated environment may optionally be an augmented reality environment including a camera view of the physical environment, or a computer-generated environment displayed on a display such that it is superimposed on a view of the physical environment which is visible through the transparent portion of the display. As shown in Figure 7M, user 7302 is standing in a physical environment (e.g., scene 105) in which a computer system (e.g., computer system 101) is operating (e.g., holding device 7100 or wearing an HMD). In some embodiments, as in the embodiments shown in Figures 7M-7R, device 7100 is a handheld device (e.g., a mobile phone, tablet, or other mobile electronic device) including a display, touch-sensitive display, etc. In some embodiments, device 7100 is a head-up display or head-mounted display This represents a wearable headset, including a playset, which can be optionally replaced. In some embodiments, the physical environment includes one or more physical surfaces and physical objects (e.g., room walls (e.g., front wall 7304, side walls 7306), floor 7308, and boxes 7502 and 7504) (e.g., a table, speaker, lamp, fixture, etc.) surrounding the user 7302. In some embodiments, one or more physical surfaces of physical objects in the environment are visible through a display generation component of a computer system (e.g., on the display of device 7100 or via the HMD).
[0156] In the embodiments shown in Figures 7M to 7R, a computer-generated three-dimensional environment corresponding to the physical environment (for example, a portion of the physical environment that is within the field of view of one or more cameras of device 7100 or visible through the transparent portion of the display of device 7100) is displayed on device 7100. The physical environment includes physical objects having representations corresponding to the computer-generated three-dimensional environment shown by the display generation component of the computer system. For example, the front wall 7304 is represented by the front wall representation 7304', the side wall 7306 is represented by the side wall representation 7306', the floor 7308 is represented by the floor representation 7308', and the boxes 7502 and 7504 are represented by the box representations 7502' and 7504' in the computer-generated environment shown on the display (for example, the computer-generated environment is an augmented reality environment that includes physical objects as part of a live view of one or more cameras of device 7100, or representations 7304', 7306', 7308', 7502', and 7504' of physical objects that are visible through the transparent portion of the display of device 7100). In some embodiments, the computer-generated environment shown on the display also includes virtual objects. According to some embodiments, as the field of view of the device 7100 with respect to the physical environment changes (for example, as the field of view of the device 7100 or one or more cameras of the device 7100 with respect to the physical environment changes in response to the movement and / or rotation of the device 7100 within the physical environment), the field of view of the computer-generated environment displayed on the device 7100 changes accordingly (including, for example, changes in the field of view of physical surfaces and physical objects (e.g., walls, floors, furniture, etc.)).
[0157] In some embodiments, if the level of interaction between user 7302 and the three-dimensional environment is below a first predetermined level (for example, if the user is simply viewing the three-dimensional environment without focusing on a specific location within it), the computer system displays the initial state of the three-dimensional environment, as shown in Figure 7N, where the representations 7502' and 7504 of boxes 7502 and 7504 are not displayed with any corresponding user interface or virtual objects.
[0158] In Figures 7O and 7P, the computer system detects that the level of interaction between the user and the three-dimensional environment has risen above a first predetermined level. In particular, in Figure 7O, eye-gaze input is detected on the representation 7502' of box 7502 (e.g., a speaker or tabletop) without simultaneous gesture input or indication that gesture input is about to be provided (e.g., the user's hands are not ready to provide gesture input). In response to detecting eye-gaze input on the representation 7502' of box 7502 in the three-dimensional environment, the computer system determines that the level of interaction between the user and box 7502 or representation 7502' has reached a first predetermined level (but not a second predetermined level that exceeds the first predetermined level). In response to determining that the level of interaction with box 7502 or representation 7502' has reached a first predetermined level, the computer system displays a first user interface 7510 corresponding to box 7502 at a location in the three-dimensional environment corresponding to the location of box 7502 in the physical environment. For example, as shown in Figure 7O, multiple user interface objects (e.g., user interface objects 7506 and 7508) are displayed so as to overlap the top surface of box 7502 or replace part of representation 7502'. In some embodiments, box 7502 is a table, and user interface objects 7506 and 7508 include one or more of the following: a virtual newspaper, a virtual screen, notifications from an application or communication channel, a keyboard and display, a sketchpad, etc. In some embodiments, box 7502 is a speaker, and user interface objects 7506 and 7508 include a volume indicator, play / pause control, the name of the song / album currently being played, today's weather, etc. In some embodiments, box 7502 is a smart lamp or device, and user interface objects 7506 and 7508 include one or more of the following: brightness or temperature control, start / stop or on / off button, timer, etc.
[0159] In Figure 7P, the gaze input moves from representation 7502' of box 7502 (e.g., a tabletop, speaker, smart lamp, or device) to representation 7504' of box 7504 (e.g., a smart medicine cabinet) without simultaneous gesture input or indication that gesture input is about to be provided (e.g., the user's hands are not ready to provide gesture input). In response to detecting that the gaze input has moved from representation 7502' of box 7502 to representation 7504' of box 7504 in the three-dimensional environment, the computer system determines that the level of interaction between the user and box 7504 or representation 7504' has reached a first predetermined level (but not a second predetermined level above the first predetermined level), or that the level of interaction between the user and box 7502 or representation 7502' has decreased to below the first predetermined level. In accordance with the determination that the level of interaction between the user and box 7502 or representation 7502' has fallen below a first predetermined level, the computer system stops displaying the first user interface 7510 corresponding to box 7502. In response to the determination that the level of interaction with box 7504 or representation 7504' has reached a first predetermined level, the computer system displays the first user interface 7512 corresponding to box 7504 at a position in the three-dimensional environment corresponding to the position of box 7504 in the physical environment. For example, as shown in Figure 7P, multiple user interface objects (e.g., user interface objects 7514 and 7516) are displayed so as to overlap the front of box 7504 or replace a part of representation 7504'. In some embodiments, box 7504 is a smart medicine shelf, and a number of user interface objects (e.g., user interface objects 7514 and 7516) include one or more of the medicine shelf's statuses (e.g., an indicator that a particular medicine or supply is running low and needs to be replenished, or a reminder that medicines for today's date have been taken).
[0160] In Figures 7Q and 7R, the computer system detects that the level of interaction between the user and the three-dimensional environment has risen above a second predetermined level, which is above a first predetermined level. In particular, in Figure 7Q, in addition to detecting gaze input on the representation 7502' of box 7502 (e.g., a speaker or tabletop), the computer system also detects an indication that gesture input is about to be provided (e.g., the user's hand is found ready to provide gesture input). In response to determining that the level of interaction between the user and box 7502 or representation 7502' has reached the second predetermined level, the computer system optionally displays a second user interface 7510', which is an extended version of the first user interface 7510 corresponding to box 7502. The second user interface 7510' corresponding to box 7502 is displayed at a location in the three-dimensional environment that corresponds to the location of box 7502 in the physical environment. For example, as shown in Figure 7Q, multiple user interface objects (e.g., user interface objects 7506, 7518, 7520, 7522, and 7524) are displayed so as to overlap the top surface of box 7502 or replace part of representation 7502'. In some embodiments, box 7502 is a table, and user interface objects 7506, 7518, 7520, 7522, and 7524 include one or more user interface objects shown in the first user interface 7510 and one or more other user interface objects not included in the first user interface 7510 (e.g., an extended display, a keyboard with additional keys not available in the first user interface 7510, a virtual desktop with application icons and a document list).In some embodiments, box 7502 is a speaker, and user interface objects 7506, 7518, 7520, 7522, and 7524 include one or more user interface objects shown in the first user interface 7510 and one or more other user interface objects not included in the first user interface 7510 (e.g., output routing control, a viewable media database, a search input field with a corresponding virtual keyboard, etc.). In some embodiments, box 7502 is a smart lamp or device, and user interface objects 7506, 7518, 7520, 7522, and 7524 include one or more user interface objects shown in the first user interface 7510 and one or more other user interface objects not included in the first user interface 7510 (e.g., various settings such as smart lamp or device, color control, scheduling control, etc.).
[0161] In some embodiments, Figure 7Q follows Figure 7O, and the second user interface 7510' is displayed in response to the user's hands being ready while the user's gaze is focused on box 7502. In some embodiments, Figure 7Q follows Figure 7P, and the user interface is displayed in response to the user's hands being ready and the user's gaze shifting from box 7504 to box 7502 (for example, the display of the first user interface 7512 stops after the gaze input leaves box 7504).
[0162] In Figure 7R, while the user's hand is ready to provide gesture input, the gaze input shifts from the representation 7502' of box 7502 (e.g., a tabletop, speaker, smart lamp, or fixture) to the representation 7504' of box 7504 (e.g., a smart medicine cabinet). In response to detecting that the gaze input has shifted from the representation 7502' of box 7502 to the representation 7504' of box 7504 in the three-dimensional environment, the computer system determines that the level of interaction between the user and box 7504 or representation 7504' has reached a second predetermined level, and that the level of interaction between the user and box 7502 or representation 7502' has decreased to below the second predetermined level and the first predetermined level. In accordance with the determination that the level of interaction between the user and box 7502 or representation 7502' has decreased to below the first predetermined level, the computer system stops displaying the second user interface 7510' corresponding to box 7502. In response to determining that the level of interaction with box 7504 or representation 7504' has reached a second predetermined level, the computer system displays a second user interface 7512' corresponding to box 7504 at a location in the three-dimensional environment corresponding to the location of box 7504 in the physical environment. For example, as shown in Figure 7R, a plurality of user interface objects (e.g., user interface objects 7514, 7516, 7526, 7528, and 7530) are displayed so as to overlap the front of box 7504 or replace a portion of representation 7504'. In some embodiments, box 7504 is a smart medicine cabinet, and the plurality of user interface objects (e.g., user interface objects 7514 and 7516) include one or more user interface objects shown in the first user interface 7510, and one or more other user interface objects not included in the first user interface 7510, such as a list of medicines or supplies in the medicine cabinet, scheduling settings for medicines of the day, and temperature and authentication settings for the medicine cabinet.
[0163] In some embodiments, Figure 7R follows Figure 7Q, and the user interface 7512 is displayed in response to the user's hands being kept in a ready position and the user's gaze shifting from box 7502 to box 7504 (for example, the display of the second user interface 7512' is stopped after the eye-tracking input leaves box 7502). In some embodiments, Figure 7R follows Figure 7P, and the second user interface 7512' is displayed in response to the user's hands being in a ready position while the user's gaze is focused on box 7504. In some embodiments, Figure 7R follows Figure 7O, and the user interface 7512' is displayed in response to the user's hands being in a ready position and the user's gaze shifting from box 7502 to box 7504 (for example, the display of the first user interface 7510 is stopped after the eye-tracking input leaves box 7502).
[0164] In some embodiments, when the computer system detects that the user's hand is floating above a physical object (e.g., box 7502 or 7504) (e.g., the distance between the user's fingers and the physical object is within a threshold distance), the computer system determines that a third level of interaction has been reached and displays a third user interface corresponding to the physical object (e.g., box 7502 or 7504) with more information and / or user interface objects than the second user interface corresponding to the physical object. In some embodiments, in response to the user's hand moving away from the physical object (e.g., the distance between the user's fingers and the physical object increases beyond a threshold distance), the third user interface shrinks back to the second user interface corresponding to the physical object.
[0165] In some embodiments, the computer system performs actions in response to touch input provided on a physical surface on a physical object (e.g., box 7502 or 7504). For example, the touch input may optionally be detected by one or more sensors, such as cameras, of the computer system, as opposed to a touch sensor on a physical surface on the physical object. In some embodiments, the location of the input on the physical surface is mapped to the location of a user interface object in a first / second / third user interface corresponding to the physical object, so that the computer system can determine which action to perform according to the location of the touch input on the physical surface.
[0166] In some embodiments, the user uses their gaze within the first / second / third user interface to select a user interface object within the first / second / third user interface that corresponds to a physical object (e.g., box 7502 or 7504). The computer performs an action corresponding to the currently selected user interface object in response to a gesture input to activate a detected user interface object while the gaze input is over the currently selected user interface object.
[0167] In some embodiments, the user optionally uses a nearby physical surface to control a physical object far from the user. For example, the user may swipe on a nearby physical surface (e.g., the back or palm of the user's hand, the arm of an armchair, a controller, etc.), and the user's gesture input is detected by one or more sensors (e.g., one or more cameras in a computer system) and used to interact with the currently displayed first / second / third user interface.
[0168] In this embodiment, in Figures 7Q and 7R, a representation 7200' of hand 7200 is displayed in a computer-generated environment. The computer-generated environment does not include a representation of the user's right hand (for example, because the right hand is not within the field of view of one or more cameras of device 7100). Furthermore, in some embodiments, for example, in the embodiments shown in Figures 7Q and 7R where device 7100 is a handheld device, the user can see parts of the surrounding physical environment apart from any representation of the physical environment displayed on device 7100. For example, parts of the user's hand are visible to the user outside the display of device 7100. In some embodiments, device 7100 in these examples can be replaced with a headset having a display (e.g., a head-mounted display) that completely blocks the user's view of the surrounding physical environment. In some such embodiments, no part of the physical environment is directly visible to the user; instead, the physical environment is visible to the user through representations of parts of the physical environment displayed by the device. In some embodiments, the user's hands are invisible to the user, either directly or via the device's display, while the device continuously or periodically monitors the current state of the user's hands to determine whether they are ready to provide gesture input. In some embodiments, the device displays an indicator indicating whether the user's hands are ready to provide input gestures, provides feedback to the user, and warns the user to adjust their hand position if they wish to provide input gestures.
[0169] Figures 7S to 7X are block diagrams illustrating methods for generating a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. Figures 7S to 7X are used to illustrate processes described later, including the process shown in Figure 12.
[0170] Figures 7T to 7X show exemplary computer-generated environments corresponding to the physical environment shown in Figure 7S. As described herein with reference to Figures 7T to 7X, according to some embodiments, the computer-generated environment may optionally be an augmented reality environment including a camera view of the physical environment, or a computer-generated environment displayed on a display such that it is superimposed on a view of the physical environment which is visible through the transparent portion of the display. As shown in Figure 7T, user 7302 is standing in a physical environment (e.g., Scene 105) that operates a computer system (e.g., computer system 101) (e.g., holding device 7100 or wearing an HMD). In some embodiments, as shown in the embodiments in Figures 7T-7X, device 7100 is a handheld device (e.g., a mobile phone, tablet, or other mobile electronic device) including a display, touch-sensitive display, etc. In some embodiments, device 7100 represents a wearable headset including a head-up display or head-mounted display, etc., and is optionally substituted. In some embodiments, the physical environment includes one or more physical surfaces and physical objects surrounding user 7302 (e.g., the walls of a room (e.g., front wall 7304, side wall 7306), floor 7308, and furniture 7310). In some embodiments, one or more physical surfaces of physical objects in the environment are visible through the display generation components of the computer system (e.g., on the display of device 7100 or via the HMD).
[0171] In the embodiments shown in Figures 7T to 7X, a computer-generated three-dimensional environment corresponding to the physical environment (for example, a portion of the physical environment that is within the field of view of one or more cameras of device 7100 or visible through the transparent portion of the display of device 7100) is displayed on device 7100. The physical environment includes physical objects having representations corresponding to the computer-generated three-dimensional environment shown by the display generation component of the computer system. For example, in a computer-generated environment shown on a display, the front wall 7304 is represented by the front wall representation 7304', the side wall 7306 is represented by the side wall representation 7306', the floor 7308 is represented by the floor representation 7308', the furniture 7310 is represented by the furniture representation 7310', and the front of the furniture 7310 7312 is represented by the front representation 7312' (for example, the computer-generated environment is an augmented reality environment that includes representations 7304', 7306', 7308', 7310', and 7312' of physical objects that are part of the live view of one or more cameras of device 7100, or physical objects that are visible through the transparent portion of the display of device 7100). In some embodiments, the computer-generated environment shown on a display also includes virtual objects (for example, a virtual object 7404 stationary in a portion of the display corresponding to a portion of the floor representation 7308' of the floor 7308). According to some embodiments, as the field of view of the device 7100 with respect to the physical environment changes (for example, as the field of view of the device 7100 or one or more cameras of the device 7100 with respect to the physical environment changes in response to the movement and / or rotation of the device 7100 within the physical environment), the field of view of the computer-generated environment displayed on the device 7100 changes accordingly (including, for example, changes in the field of view of physical surfaces and physical objects (e.g., walls, floors, furniture, etc.)).
[0172] In Figure 7T, the computer system detects input corresponding to a request to activate a voice-based virtual assistant. For example, the user provides the computer system with a voice-based activation command, "Assistant!". In some embodiments, the user optionally turns to look at a predetermined location in the three-dimensional environment corresponding to the home position of the voice-based virtual assistant, and / or provides activation input (e.g., tap input with the user's finger or controller, gaze input, etc.).
[0173] In Figures 7U and 7W, in response to detecting input corresponding to a request to activate a voice-based virtual assistant in a three-dimensional environment, the computer system displays a visual representation of the virtual assistant in the three-dimensional environment. In some embodiments, the visual representation of the virtual assistant is a virtual object 7602. For example, the virtual object 7602 is the virtual assistant's avatar (e.g., a glowing ellipse or animated character). In some embodiments, the visual indication is not necessarily an object with a virtual surface, but a visual effect such as lighting around the peripheral area of the display, the peripheral area of the user's field of view, or the peripheral area of the target area of the eye-tracking input. In some embodiments, other visual effects (e.g., darkening or blurring the background of the virtual assistant or the entire display) are displayed in conjunction with the display of the virtual assistant's visual indication.
[0174] As shown in Figures 7U and 7W, the visual representation of the virtual assistant has a first set of values for a first display characteristic (e.g., luminance, color) of the visual representation when the virtual assistant is activated. For example, the visual representation is an luminescent ellipse having a first luminance value distribution and a first color value distribution across various parts of the visual representation. The computer system modifies the visual appearance of a first physical surface 7312 of a physical object 7310 in a three-dimensional environment, or its representation 7312', and the visual appearance of a first virtual surface of a virtual object 7404 in a three-dimensional environment, according to the first set of values for the first display characteristic. For example, as shown in Figures 7U and 7W, the computer system generates simulated illumination at a location in the three-dimensional environment mapped to the surface of the physical object 7310 in the physical world, and the values of the first display characteristic of the illumination take into account the spatial relationship between the virtual object 7602 and its representation 7310' in the three-dimensional world, the surface characteristics of the physical object 7310, and the simulated physical light propagation principle. In Figures 7U and 7W, the front representation 7312' of the furniture representation 7310' appears to be illuminated by the virtual object 7602, and because the virtual object 7602 is closer to the left side of the front representation 7312' in the three-dimensional environment, the left side of the front representation 7312' appears to be more strongly illuminated (for example, with higher luminance and higher color saturation from the virtual object 7602) than the right side of the front representation 7312'. Similarly, as shown in Figures 7U and 7W, the computer system generates simulated illumination at a location in the three-dimensional environment, which is mapped to the surface of the virtual object 7404 in the three-dimensional environment, and the values of the first display characteristics of the illumination take into account the spatial relationship between the virtual object 7602 and the virtual object 7404 in the three-dimensional world, the surface characteristics of the virtual object 7404, and the simulated physical light propagation principle.In Figures 7U and 7W, the top surface of virtual object 7404 appears to be illuminated by the visual representation of the virtual assistant (e.g., virtual object 7602), and because the visual representation of the virtual assistant (e.g., virtual object 7602) is closer to the top than the middle portion of the surface of virtual object 7404 in the three-dimensional environment, the middle region of the surface of virtual object 7404 appears to be illuminated less brightly than the top of virtual object 7404 (e.g., with lower luminance and lower color saturation from virtual object 7602).
[0175] In some embodiments, as shown in Figures 7U and 7W, the computer system also generates simulated shadows for physical and virtual objects under the illumination of the visual representation of the virtual assistant. For example, the computer generates a shadow 7606 behind the furniture representation 7310' in the three-dimensional environment based on the spatial relationship between the virtual object 7602 and the furniture representation 7310' in the three-dimensional environment, the surface properties of the furniture 7310, and the simulated physical light propagation principle. The computer also generates a shadow 7604 under the virtual object 7404 in the three-dimensional world based on the spatial relationship between the virtual object 7602 and the virtual object 7404 in the three-dimensional environment, the simulated surface properties of the virtual object 7404, and the simulated physical light propagation principle.
[0176] As shown in Figure 7V following Figure 7U, in some embodiments, the position of the visual indication of a virtual assistant (e.g., virtual object 7602) is fixed relative to the user's head (e.g., represented by a display (e.g., a touch-sensitive display) or (e.g., an HMD) and moves relative to the three-dimensional environment as the display moves relative to the physical world, or as the user's head (or HMD) moves relative to the physical world. In Figure 7V, as the user's head (e.g., the three-dimensional environment is shown via an HMD) or display (e.g., a touch-sensitive display) moves within the physical environment, the spatial relationship between the visual representation of the virtual assistant (e.g., virtual object 7602) in the three-dimensional environment and the representations of physical objects (e.g., furniture representation 7310') and virtual objects (e.g., virtual object 7404) changes in response to the movement, and thus the simulated lighting on the physical and virtual objects in the three-dimensional environment is adjusted. For example, since the visual representation (virtual object 7602) is now closer to the front representation 7312' than before it was moved, the front representation 7312' of the furniture representation 7310' is brightly lit (with higher brightness and color saturation than the visual representation of the virtual assistant (e.g., virtual object 7602)). Correspondingly, since the visual representation (e.g., virtual object 7602) is now further away from the virtual object 7404 than before it was moved, the top surface of the virtual object 7404 is dimly lit (with lower brightness and color saturation than the visual representation of the virtual assistant (e.g., virtual object 7602)).
[0177] In contrast to the embodiment shown in Figure 7V, in some embodiments, the position of the visual indication of the virtual assistant (e.g., virtual object 7602) is fixed relative to the three-dimensional environment, rather than to the user's head (e.g., represented by a display (e.g., a touch-sensitive display) or HMD). Therefore, the spatial relationship between the visual representation of the virtual assistant (e.g., visual representation 7602) and the physical objects (e.g., furniture representation 7310') and virtual objects (e.g., virtual object 7404) represented in the three-dimensional environment does not change when the display moves relative to the physical world or when the user moves their head (or HMD) relative to the physical world. In Figure 7V, as the user's head (e.g., the three-dimensional environment is shown via the HMD) or the display (e.g., a touch-sensitive display) moves within the physical environment, the spatial relationship between the visual representation of the virtual assistant in the three-dimensional environment and the representations of the physical and virtual objects does not change in response to the movement, and therefore the simulated lighting on the physical and virtual objects in the three-dimensional environment does not change. However, the field of view of the three-dimensional world shown on the display changes due to the movement.
[0178] In some embodiments, the various examples and embodiments described herein optionally include discrete small-motion gestures performed by moving one or more of the user's fingers relative to other fingers or parts of the user's hand to perform an action immediately before or during a gesture in order to interact with a virtual or mixed reality environment.
[0179] In some embodiments, input gestures are detected by analyzing data and signals captured by a sensor system (e.g., sensor 190 in Figure 1, image sensor 314 in Figure 3). In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras such as a motion RGB camera, an infrared camera, a depth camera, etc.). For example, one or more imaging sensors are components of a computer system (e.g., computer system 101 in Figure 1 (e.g., portable electronic device 7100 or HMD)) that includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4 (e.g., a touchscreen display that functions as a display and a touch-sensing surface, a stereoscopic display, a display with a pass-through portion, etc.)) or provide data to the computer system. In some embodiments, one or more imaging sensors include one or more rear cameras on the side of the device opposite to the device's display. In some embodiments, input gestures are detected by a sensor system of a head-mounted system (e.g., a VR headset that includes a stereoscopic display that provides a left image for the user's left eye and a right image for the user's right eye). For example, one or more cameras, which are components of a head-mounted system, are mounted on the front and / or bottom of the head-mounted system. In some embodiments, one or more imaging sensors are positioned in the space in which the head-mounted system is used (e.g., arranged around the head-mounted system at various positions in a room) so that the imaging sensors capture images of the head-mounted system and / or the user of the head-mounted system. In some embodiments, input gestures are detected by a sensor system of a head-up device (e.g., a head-up display, a car windshield capable of displaying graphics, a window capable of displaying graphics, or a lens capable of displaying graphics). For example, one or more imaging sensors are mounted on the interior of a car. In some embodiments, the sensor system includes one or more depth sensors (e.g., a sensor array).For example, one or more depth sensors include one or more light-based (e.g., infrared) sensors and / or one or more acoustic-based (e.g., ultrasonic) sensors. In some embodiments, the sensor system includes one or more signal emitters, such as light emitters (e.g., infrared emitters) and / or sound emitters (e.g., ultrasonic emitters). For example, while light (e.g., light from an infrared light emitter array having a predetermined pattern) is projected onto a hand (e.g., hand 7200), an image of the hand under illumination is captured by one or more cameras, and the captured image is analyzed to determine the position and / or configuration of the hand. By determining input gestures using signals from image sensors directed at the hand, in contrast to using signals from a touch-sensing surface or other direct contact or proximity-based mechanisms, the user is free to choose whether to perform large movements or remain relatively stationary when providing input gestures with their hand, without experiencing constraints imposed by a particular input device or input area.
[0180] In some embodiments, a microtap input indicates a tap input of the thumb on the index finger of the user's hand (e.g., on the side of the index finger adjacent to the thumb). In some embodiments, the tap input is detected without the need to lift the thumb from the side of the index finger. In some embodiments, the tap input is detected according to the determination that a downward movement of the thumb is followed by an upward movement of the thumb, and the thumb is in contact with the side of the index finger for less than a threshold time. In some embodiments, a tap-hold input is detected according to the determination that the thumb moves from an elevated position to a touch-down position and remains in the touch-down position for at least a first threshold time (e.g., a tap-time threshold or another time threshold longer than the tap-time threshold). In some embodiments, the computer system requires the entire hand to remain substantially stationary in a position for at least a first threshold time in order to detect a tap-hold input by the thumb on the index finger. In some embodiments, a touch-hold input is detected without requiring the hand to remain substantially stationary (e.g., the entire hand can move while the thumb is placed on the side of the index finger). In some embodiments, tap-hole drag input is detected when the thumb touches the side of the index finger and the entire hand moves while the thumb remains stationary on the side of the index finger.
[0181] In some embodiments, a microflick gesture indicates a push or flick input of the thumb moving across the index finger (e.g., from the palmar side to the rear side of the index finger). In some embodiments, an extension of the thumb involves an upward movement away from the side of the index finger, such as an upward flick input by the thumb. In some embodiments, the index finger moves in the opposite direction to the thumb while the thumb moves forward and upward. In some embodiments, a reverse flick input is performed by the thumb moving from an extended position to a retracted position. In some embodiments, the index finger moves in the opposite direction to the thumb while the thumb moves backward and downward.
[0182] In some embodiments, the micro-swipe gesture is a swipe input by moving the thumb along the index finger (e.g., along the side of the index finger adjacent to the thumb or along the side of the palm). In some embodiments, the index finger is optionally extended (e.g., substantially straight) or flexed. In some embodiments, the index finger moves between the extended and flexed states while the thumb moves in the swipe input gesture.
[0183] In some embodiments, different phalanges of different fingers correspond to different inputs. Microtap inputs of the thumb across different phalanges of different fingers (e.g., index finger, middle finger, ring finger, and optionally little finger) are optionally mapped to different actions. Similarly, in some embodiments, different push or click inputs performed by the thumb across different fingers and / or different parts of the fingers can trigger different actions at each user interface contact. Likewise, in some embodiments, different swipe inputs performed by the thumb along different fingers and / or in different directions (e.g., towards the distal or proximal end of the finger) trigger different actions in each user interface context.
[0184] In some embodiments, the computer system processes tap input, flick input, and swipe input as different types of input based on the type of thumb movement. In some embodiments, the computer system processes input having different finger positions tapped, touched, or swiped by the thumb as different sub-input types (e.g., proximal, intermediate, distal subtypes, or index finger, middle finger, ring finger, or little finger subtypes) of a given input type (e.g., tap input type, flick input type, swipe input type, etc.). In some embodiments, the amount of movement performed by the moving finger (e.g., thumb), and / or other measures of movement associated with the finger movement (e.g., speed, initial speed, ending speed, duration, direction, movement pattern, etc.) are used to quantitatively influence the action triggered by the finger input.
[0185] In some embodiments, the computer system recognizes combination input types that combine a series of thumb movements, such as tap-swipe input (e.g., the thumb swiping along the side of a finger after touching down another finger), tap-flick input (e.g., the thumb flicking across a finger from the side of the palm to the back of the finger after touching down another finger), and double-tap input (e.g., two consecutive taps on the side of a finger in approximately the same position).
[0186] In some embodiments, gesture input is performed by the index finger instead of the thumb (e.g., the index finger performs a tap or swipe on the thumb, or the thumb and index finger move toward each other to perform a pinch gesture). In some embodiments, wrist movement (e.g., a wrist flick in a horizontal or vertical direction) is performed immediately before, immediately after (e.g., within a threshold time), or concurrently with finger movement input to trigger an additional, different, or modified action in the current user interface context compared to finger movement input without modification by wrist movement. In some embodiments, finger input gestures performed with the user's palm facing the user's face are treated as a different type of gesture than finger input gestures performed with the user's palm facing away from the user's face. For example, a tap gesture performed with the user's palm facing the user performs an action with added (or reduced) privacy protection compared to an action performed in response to a tap gesture performed with the user's palm facing away from the user's face (e.g., the same action).
[0187] In the embodiments provided herein, one type of finger input can be used to trigger an action type, but in other embodiments, other types of finger input may be optionally used to trigger the same type of action.
[0188] Further explanations regarding Figures 7A to 7X are provided below with reference to methods 8000, 9000, 10000, 11000, and 12000 described with respect to Figures 8 to 12 below.
[0189] Figure 8 is a flowchart of Method 8000 for interacting with a computer-generated three-dimensional environment (including, for example, reconstruction and other interactions) according to several embodiments. In some embodiments, Method 8000 is executed on a computer system (e.g., computer system 101 in Figure 1) which includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more input devices (e.g., one or more cameras (e.g., cameras facing downward at the user's hands or facing forward from the user's head (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras)), a controller, a touch-sensing surface, a joystick, a button, etc.). In some embodiments, Method 8000 is executed by instructions stored on a non-temporary computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., control unit 110 in Figure 1A). Some operations of method 8000 are arbitrarily combined, and / or the order of some operations is arbitrarily changed.
[0190] In method 8000, the computer system displays a virtual object (e.g., virtual objects 7208(a-1) and 7B(a-1) in Figure 7A) at a first spatial location in a three-dimensional environment (e.g., a physical environment visible through a display generation component, an imitated reality environment, a virtual reality environment, an augmented reality environment, a mixed reality environment, etc.) (8002). While the virtual object (e.g., virtual object 7208) is displayed at the first spatial location in the three-dimensional environment, the computer system detects a first hand movement performed by the user (8004) (e.g., detecting a movement of the user's fingers and / or wrist that satisfies one or more gesture recognition criteria). In response to detecting a first hand movement performed by the user (8006), and according to the determination that the first hand movement satisfies a first gesture criterion (for example, the first hand movement is a pinch and drag gesture (for example, a pinch of fingers resulting from moving the entire hand laterally) or a swipe gesture (for example, a micro-swipe gesture with a finger across the surface of another finger or a controller)), the computer system performs a first action according to the first hand movement without moving the virtual object from a first spatial position (for example, a pinch and drag gesture before entering reconstruction mode does not move the object from one position to another) (for example, rotating the virtual object, adjusting controls associated with the virtual object, navigating the virtual object (turning the pages of a virtual book), etc.). This is shown, for example, in Figures 7A(a-1) to 7A(a-3), and Figure 7A(a-1), followed by Figures 7A(a-4) and 7A(a-5).In response to detecting a first hand movement performed by the user (8006), and according to the determination that the first hand movement satisfies a second gesture criterion (e.g., a pinch gesture and a wrist flick gesture following the pinch gesture (e.g., the pinch movement of the fingers results from rotating the hand around the wrist (e.g., flicking upward or to the side))), the computer system displays a first visual indication that the virtual object has entered reconstruction mode (e.g., the device activates the reconstruction mode of the virtual object, the virtual object is removed from its original position and / or becomes semi-transparent and floats above its original position). This is shown, for example, in Figures 7B(a-1) to 7B(a-3). While the computer displays the virtual object with the first visual indication that the virtual object has entered reconstruction mode, The data system detects a second hand movement performed by the user (8008). In response to the detection of a second hand movement performed by the user, and according to the determination that the second hand movement satisfies a first gesture criterion (for example, the first hand movement is a pinch and drag gesture (for example, a pinch finger movement results from the entire hand moving laterally)), the computer system moves the virtual object from a first spatial position to a second spatial position (for example, without performing the first action) according to the second hand movement (8010) ((for example, once in reconfiguration mode, a wrist flick no longer needs to be continued, and a simple pinch and drag gesture moves the object from one position to another)). This is shown, for example, in Figures 7B(a-3) to 7B(a-6) following Figure 7B(a-2), or in Figures 7B(a-5) and 7B(a-6).
[0191] In some embodiments, in method 8000, in response to detecting a first hand movement performed by the user, and according to a determination that the first hand movement satisfies a third gesture criterion (for example, the first hand movement is a microtap gesture without lateral or rotational movement of the entire hand), the computer system performs a second action corresponding to the virtual object (for example, activating a function corresponding to the virtual object (e.g., launching an application, establishing a communication session, displaying content, etc.)). In some embodiments, in response to detecting a second hand movement performed by the user, and according to a determination that the second hand movement satisfies a third gesture criterion (for example, the second hand movement is a microtap gesture without lateral or rotational movement of the entire hand), the device stops displaying a first visual indication that the virtual object has entered reconstruction mode, indicating that the virtual object has exited reconstruction mode (for example, the device deactivates the virtual object's reconstruction mode and, if the virtual object has not been moved, returns the virtual object to its original position, or, if the virtual object has been moved by user input, places the virtual object in a new position to restore the virtual object's original appearance). In some embodiments, in response to detecting a second hand movement performed by the user, and according to a determination that the second hand movement does not satisfy a first gesture criterion (for example, the second hand movement is a free hand movement that does not satisfy pinching the fingers together or another predetermined gesture criterion), the device maintains the virtual object in reconfiguration mode without moving the virtual object. In other words, while the virtual object is in reconfiguration mode, the user can move their hand in a manner that does not correspond to a gesture that moves the virtual object and does not cause the virtual object to exit reconfiguration mode. For example, the user can use this opportunity to explore the three-dimensional environment and then prepare a suitable position to move the virtual object.
[0192] In some embodiments, the second hand movement does not satisfy the second gesture criterion (for example, the second hand movement is not a pinch gesture and a wrist flick gesture following a pinch gesture (for example, a finger pinch movement results from rotating the hand around the wrist (e.g., flicking upward or sideways))).
[0193] In some embodiments, the second gesture criterion includes requirements that are met by a pinch gesture and a wrist flick gesture following the pinch gesture (for example, the second gesture criterion is met with respect to the virtual object when the thumb and index finger of the hand move to a position in three-dimensional space corresponding to the position of the virtual object and come into contact with each other, and then the entire hand rotates around the wrist while the thumb and index finger remain in contact with each other).
[0194] In some embodiments, the second gesture criterion includes a requirement that be satisfied by a wrist flick gesture detected while the object selection criterion is being met (for example, the second gesture criterion is met for a virtual object when the entire hand rapidly rotates around the wrist while the virtual object is currently selected (for example, by a previous selection input (e.g., gaze input directed at the virtual object, a pinch gesture directed at the virtual object, a two-finger tap gesture directed at the virtual object, etc.)). In some embodiments, the previous selection input may be in progress (e.g., a pinch gesture or gaze input) or completed (e.g., a two-finger tap gesture to select the virtual object) when the wrist flick gesture is detected.
[0195] In some embodiments, the first gesture criterion includes requirements that are satisfied by a movement input provided by one or more fingers of a hand (e.g., a single or multiple fingers moving laterally together) (e.g., a lateral movement of a finger in the air or across a surface (e.g., the surface of a controller or the surface of another finger), or a tap movement of a finger in the air or on a surface (e.g., the surface of a controller or the surface of another finger)).
[0196] In some embodiments, while displaying a virtual object with a first visual indication that the virtual object has entered reconstruction mode, the computer system detects a predetermined input specifying the target location of the virtual object in a three-dimensional environment (for example, detecting a predetermined input includes detecting a shift in the user's gaze from a first spatial location to a second spatial location, or detecting a tap input by a finger (on a controller or a tap in the air or on a surface with the same hand) while the user's gaze is focused on the second spatial location in three-dimensional space). In response to detecting the predetermined input specifying the target location of the virtual object in the three-dimensional environment, the computer system displays a second visual indication (for example, a glowing or shadowed overlay (of the shape of the virtual object)) at the target location before moving the virtual object from the first spatial location to the target location (for example, the second spatial location or a location different from the second spatial location). In some embodiments, the second visual indication is displayed at the target location in response to detecting a predetermined input before detecting a second hand movement that actually moves the virtual object. In some embodiments, a second hand movement that satisfies the first gesture criterion is detected after the target position of the virtual object is specified by a predetermined input (e.g., gaze input, tap input) provided while the virtual object is in reconstruction mode, and may be a tap input, finger flick input, hand swipe input, or pinch and drag input. In some embodiments, the predetermined input is detected before the second hand movement is detected (e.g., if the predetermined input is a gaze input or tap input that selects the target position of the virtual object (e.g., the user may look away from the target position after providing the predetermined input), the second hand movement is a small finger flick or finger tap without movement of the entire hand that initiates the movement of the virtual object toward the target position).In some embodiments, a predetermined input is detected simultaneously with a second hand movement (for example, if the predetermined input is a gaze input focused on the target position of a virtual object (for example, the user maintains their gaze on the target position while the second movement (e.g., a small finger flick or tap without moving the entire hand) initiates the movement of the virtual object toward the target position)). In some embodiments, the predetermined input is a second hand movement (for example, the predetermined input is a pinch gesture that grasps the virtual object and drags it toward the target position).
[0197] In some embodiments, detecting a predetermined input that specifies the target position of a virtual object in a three-dimensional environment includes detecting movement in the predetermined input (e.g., movement of gaze input or movement of a finger before a tap), and displaying a second visual indication (e.g., a glowing or shadowed overlay (e.g., the shape of the virtual object)) at the target position includes updating the position of the second visual indication based on the movement of the predetermined input (e.g., the position of the glowing or shadowed overlay (e.g., the shape of the virtual object) is changed continuously and dynamically in accordance with the movement of the gaze input and / or the position of the finger before a tap of the input).
[0198] In some embodiments, after the completion of a second hand movement that satisfies the first gesture criterion, and while the virtual object remains in reconstruction mode (for example, after the object has been moved in accordance with the second hand movement, and while the virtual object is displayed in a first visual indication that the virtual object has entered reconstruction mode), the computer system detects a third hand movement that satisfies the first gesture criterion (for example, a micro-swipe gesture in which the thumb swipes across the side of the index finger of the same hand, or a swipe gesture by a finger on the touch-sensitive surface of a controller). In response to the detection of the third hand movement, the computer system moves the virtual object from its current position to a third spatial position in accordance with the third hand movement.
[0199] In some embodiments, the three-dimensional environment includes one or more planes (e.g., the surface of a physical object, a surface that simulates a virtual object, the surface of a virtual object representing a physical object, etc.), and moving a virtual object from a first spatial position to a second spatial position in accordance with a second hand movement includes constraining the movement path of the virtual object to the first plane of one or more planes during the movement of the virtual object in accordance with the second hand movement (for example, if the first and second spatial positions lie on the same plane, the virtual object slides along the plane even if the movement path of the second hand movement does not strictly follow the plane).
[0200] The method according to any one of claims 1 to 10. In some embodiments, the three-dimensional environment includes at least a first plane and a second plane (e.g., the surface of a physical object, a surface that mimics a virtual object, the surface of a virtual object representing a physical object, etc.), and moving a virtual object from a first spatial position to a second spatial position in accordance with a second hand movement includes constraining the path of the virtual object to the first plane during the first part of the movement of the virtual object in accordance with the second hand movement, constraining the path of the virtual object to the second plane during the second part of the movement of the virtual object in accordance with the second hand movement, and raising the altitude of the virtual object during the third part of the movement of the virtual object between the first and second parts of the movement of the virtual object (e.g., the object jumps when switching between planes of the real world).
[0201] In some embodiments, in response to detecting a first hand movement performed by the user, and according to a determination that the first hand movement satisfies a second gesture criterion (e.g., a pinch gesture followed by a wrist flick gesture (e.g., a finger pinch movement resulting from a hand rotating around the wrist (e.g., flicking upward or sideways))), the computer system generates an audio output in conjunction with displaying a first visual indication that the virtual object has entered reconstruction mode (e.g., the device generates a discrete audio output (e.g., a beep or chirp) that provides an indication that the virtual object has been removed from its original position, and / or generates a continuous audio output (e.g., continuous music or sound waves) while the virtual object remains in reconstruction mode).
[0202] In some embodiments, while the virtual object is in reconstruction mode, the computer system detects a second hand movement, moves the virtual object according to the second movement, and then detects a fourth hand movement. In response to the detection of the fourth hand movement, and according to the determination that the fourth hand movement satisfies a first gesture criterion, the computer system moves the virtual object from a second spatial position to a third spatial position according to the fourth hand movement, and according to the determination that the fourth hand movement satisfies a fourth gesture criterion (e.g., a pinch gesture followed by a wrist flick gesture (e.g., a finger pinch movement resulting from a hand rotating around the wrist (e.g., a downward flick))), the computer system stops displaying the first visual indication to indicate that the virtual object has exited reconstruction mode. In some embodiments, the device displays an animation of the virtual object being positioned at a third spatial position in a three-dimensional environment, in conjunction with stopping the display of the first visual indication (e.g., restoring the normal appearance of the virtual object).
[0203] It should be understood that the specific order of operations described in Figure 8 is merely an example, and is not intended to indicate that the described order is the only order in which the operations can be performed. Those skilled in the art will recognize various methods for reordering the operations described herein. In addition, it should be noted that details of other processes described herein with respect to other methods described herein (e.g., methods 9000, 10000, 11000, and 12000) are also applicable in a manner similar to method 8000 described above in relation to Figure 8. For example, the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to method 8000 optionally have one or more characteristics of the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to other methods described herein (e.g., methods 9000, 10000, 11000, and 12000). For brevity, those details will not be repeated here.
[0204] Figure 9 is a flowchart of Method 9000, according to several embodiments, for generating a computer-generated three-dimensional environment (including, for example, simulating visual interaction between physical and virtual objects). In some embodiments, Method 9000 is executed on a computer system (e.g., computer system 101 in Figure 1) which includes display generation components (e.g., display generation components 120 in Figures 1, 3, and 4) (e.g., head-up displays, displays, touchscreens, projectors, etc.) and one or more input devices (e.g., cameras (e.g., cameras facing downwards at the user's hands or forwards from the user's head (e.g., color sensors, infrared sensors, and other depth-sensing cameras)), controllers, touch-sensing surfaces, joysticks, buttons, etc.). In some embodiments, Method 9000 is executed by instructions stored in a non-temporary computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., control unit 110 in Figure 1A). Some operations of method 9000 are arbitrarily combined, and / or the order of some operations is arbitrarily changed.
[0205] In method 9000, the computer system displays a three-dimensional scene via a display generation component, which includes at least a first virtual object (e.g., virtual object 7332 in Figures 7E and 7F) at a first location (e.g., a virtual window on a wall, a virtual screen on a wall displaying a video) and a first physical surface (e.g., a front wall 7304, a side wall 7306, a floor 7308, furniture 7310, or a representation thereof) at a second location separate from the first location (e.g., a bookshelf in a room away from a wall, a wall, or the floor of a room) (for example, the first virtual object and the first physical surface are separated by real or simulated free space) (9002), and the virtual object is virtual The object is displayed with first values relating to first display characteristics corresponding to first parts of the object (e.g., luminance and color values in first parts 7332-b and 7332-c, and 7332-b' and 7332-c') and second values relating to first display characteristics corresponding to second parts of the virtual object (e.g., luminance and color values in second parts 7332-a and 7332-d, and 7332-a' and 7332-d') (for example, the virtual object has different luminance values or colors in different parts of the virtual object, and the first display characteristics are independent of the shape or dimensions of the virtual object), and the second values relating to first display characteristics are different from the first values relating to first display characteristics. While displaying a three-dimensional scene including a first virtual object and a first physical surface, the computer system generates a first visual effect at a second location in the three-dimensional scene (e.g., the location of a physical surface in the scene) via a display generation component (9004).Generating a first visual effect includes modifying the visual appearance of a first part of a first physical surface in a three-dimensional scene according to a first value of a first display characteristic corresponding to a first part of a first virtual object, and modifying the visual appearance of a second part of a first physical surface in a three-dimensional scene according to a second value of a first display characteristic corresponding to a second part of a first virtual object, wherein the visual appearance of the first part of the first physical surface and the visual appearance of the second part of the first physical surface are modified differently due to the difference between the first and second values of the first display characteristic in the first and second parts of the first virtual object (for example, different color and brightness values of different parts of the virtual object cause different color and brightness values of different parts of the physical surface due to the spatial relationship between different parts of the virtual object and different parts of the physical surface) (for example, according to the simulated spatial relationship between the virtual object and the physical surface, the actual and simulated physical properties of the virtual object and the physical surface, and the simulated physical principles). This is shown, for example, in Figures 7E-7F.
[0206] In some embodiments, the computer system detects changes in the appearance of a first virtual object, including changes in the values of first display characteristics in first and second parts of the first virtual object. In response to detecting changes in the appearance of the first virtual object, the computer system modifies the visual appearance of the first physical surface in various parts of the first physical surface in accordance with the changes in the appearance of the first virtual object. The modification includes modifying the visual appearance of a first part of the first physical surface according to a first relationship between the first display characteristics and the visual appearance of the first part of the first physical surface, and modifying the visual appearance of a second part of the first physical surface according to a second relationship between the first display characteristics and the visual appearance of the second part of the first virtual object, where the first and second relationships correspond to different physical characteristics of the first and second parts of the first physical surface. For example, the first and second relationships are based on the pseudo-physical laws of light emitted from a virtual object interacting with the first physical surface, but differ by the distance, shape, surface texture, and optical properties corresponding to various parts of the first physical surface, and / or the different spatial relationships between the various parts of the first physical surface and each corresponding part of the first virtual object.
[0207] In some embodiments, the first virtual object includes a virtual overlay (e.g., a virtual window showing a virtual landscape (e.g., seen through a window) on a second physical surface (e.g., a wall) at a location corresponding to a first location in a three-dimensional scene) (the first virtual object is a virtual window displayed at a location corresponding to a physical window or part of a physical wall in the real world), and the computer system changes the appearance of the virtual overlay (e.g., changes the appearance of the landscape shown in the virtual overlay) according to changes in the values of one or more parameters, including at least one of time, location, and size of the virtual overlay. For example, as time changes in the real world or in a user-defined setting, the device changes the virtual landscape (e.g., a view of a city, nature, landscape, factory, etc.) shown in the virtual overlay (e.g., a virtual window) according to the change in time. In another embodiment, the user or device specifies the scene location of the virtual landscape shown in the virtual overlay, and the virtual landscape is selected from a landscape database based on the scene location. In another embodiment, the user requests the computer system to increase or decrease the size of the virtual overlay (for example, moving from a small virtual window to a large virtual window, or replacing an entire wall with a virtual window), and the computer system changes the amount of virtual scenery presented through the virtual overlay.
[0208] In some embodiments, generating a first visual effect includes modifying the visual appearance of a first portion of a first physical surface (e.g., an opposing wall or floor in the real world) in accordance with changes in the content shown in a first portion of a virtual overlay, and modifying the visual appearance of a second portion of the first physical surface in accordance with changes in the content shown in a second portion of a virtual overlay. For example, on a real-world floor, the amount, color, and direction of light coming from different portions of a virtual window superimposed on a physical wall (e.g., depending on the time of day) creates different simulated illumination on the floor in front of the virtual window. A computer system generates a second virtual overlay for the floor that simulates the different amounts, colors, and directions of illumination in different portions of a second virtual overlay corresponding to different portions of the floor. For example, as the time of day changes, the amount and direction of light corresponding to the virtual window change, and accordingly, the amount of simulated illumination shown in the second virtual overlay on the floor also changes (e.g., the direction of light, color, and hue are different in the morning, noon, and evening).
[0209] In some embodiments, the first virtual object includes a virtual screen at a position corresponding to a first position in a three-dimensional scene that displays media content (e.g., a flat virtual screen for displaying video or moving images, a three-dimensional space or dome surface for displaying three-dimensional video or an immersive holographic experience from the user's viewpoint) (for example, the virtual screen is freestanding and not attached to any physical surface, or superimposed on a physical surface such as a wall or television screen), and the computer system changes the content displayed on the virtual screen as the playback of the media item progresses. For example, as video or moving image playback progresses, the content is displayed on the virtual screen (e.g., 2D or 3D or immersive) according to the current playback position of the video or moving image.
[0210] The method according to claim 18. In some embodiments, generating a first visual effect includes modifying the visual appearance of a first portion of a first physical surface (e.g., an opposing wall or floor in the real world) in accordance with a change in content shown in a first portion of a virtual screen, and modifying the visual appearance of a second portion of the first physical surface in accordance with a change in content shown in a second portion of a virtual screen. For example, on the surface of physical objects in the surrounding environment (e.g., floor, wall, couch, user's body, etc.), the amount, color, and direction of light coming from different portions of the virtual screen produce different simulated illumination on the surface of physical objects in the surrounding environment. The device generates a virtual overlay on the surrounding physical surface that simulates different amounts, colors, and directions of illumination in different portions of the virtual overlay corresponding to different portions of the physical surface. As the video scene changes, the amount, color, and direction of light also change, altering the simulated illumination superimposed on the surrounding physical surface.
[0211] In some embodiments, the first virtual object is a virtual assistant that interacts with the user through speech (for example, the virtual assistant is activated in various contexts and provides assistance to the user with respect to various tasks and interactions with electronic devices), and the computer system changes the appearance of the virtual assistant according to the virtual assistant's operating mode. For example, when the virtual assistant is performing various tasks or operating in various modes (e.g., standby, listening to user commands, moving from one location to another, performing a task according to user commands, completing a task, performing various types of tasks, etc.), the color, size, hue, brightness, etc. of the virtual assistant change. As a result of the changes in the appearance of the virtual assistant, the device generates simulated lighting on the physical surface at positions corresponding to the positions surrounding the virtual assistant.
[0212] In some embodiments, generating a first visual effect involves modifying the visual appearance of a first portion of a first physical surface according to the imitation reflection of a first virtual object on a first portion of the first physical surface (for example, the imitation reflection is generated according to the surface properties of the first portion of the first physical surface, the relative positions of the first virtual object and the first portion of the first physical surface in a three-dimensional scene, the imitation physical properties of the light emitted from the first virtual object, and the physical light propagation principles that determine how light is reflected and transmitted and how an object is illuminated by real-world light). In some embodiments, generating a first visual effect further includes modifying the visual appearance of the second portion of the first physical surface according to the imitation reflection of the first virtual object on the second portion of the first physical surface (for example, the imitation reflection is generated according to the surface properties of the second portion of the first physical surface, the relative positions of the first virtual object and the second portion of the first physical surface in a three-dimensional scene, the imitation physical properties of the light emitted from the first virtual object, and the physical light propagation principles that determine how light is reflected and transmitted and how the object is illuminated by real-world light).
[0213] In some embodiments, generating a first visual effect involves altering the visual appearance of a first portion of a first physical surface (e.g., a non-reflective physical surface) according to a simulated shadow cast by a first virtual object on a first portion of a first physical surface (e.g., the simulated shadow is generated by the device according to the surface properties of the first portion of the first physical surface, the relative positions of the first virtual object and the first portion of the first physical surface in a three-dimensional scene, the simulated physical properties of the first virtual object (e.g., shape, size, etc.), the actual light source, a simulated light source present in the three-dimensional scene, and the principles of physical light propagation and refraction). In some embodiments, generating a first visual effect further includes modifying the visual appearance of the second portion of the first physical surface (e.g., a non-reflective physical surface) according to a simulated shadow of the first virtual object on the second portion of the first physical surface (e.g., the simulated shadow is generated by the device according to the surface properties of the second portion of the first physical surface, the relative position of the first virtual object and the second portion of the first physical surface in a three-dimensional scene, the simulated physical properties of the first virtual object (e.g., shape, size, etc.), actual light sources, simulated light sources present in the three-dimensional scene, and the physical light propagation principles that determine how light is reflected and transmitted and how objects are illuminated by this light in the real world).
[0214] It should be understood that the specific order of operations described in Figure 9 is merely an example, and is not intended to indicate that the described order is the only order in which the operations can be performed. Those skilled in the art will recognize various methods for rearranging the operations described herein. In addition, it should be noted that details of other processes described herein with respect to other methods described herein (e.g., methods 8000, 10000, 11000, and 12000) are also applicable in a manner similar to method 9000 described above in relation to Figure 9. For example, the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to method 9000 optionally have one or more characteristics of the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to other methods described herein (e.g., methods 8000, 10000, 11000, and 12000). For brevity, those details will not be repeated here.
[0215] Figure 10 is a flowchart of Method 10000, in several embodiments, for generating a computer-generated three-dimensional environment and facilitating user interaction with the three-dimensional environment (including, for example, gradually adjusting the level of the computer-generated experience based on user input). In some embodiments, Method 10000 is executed in a computer system (e.g., computer system 101 in Figure 1) which includes display generation components (e.g., display generation components 120 in Figures 1, 3, and 4) (e.g., head-up displays, displays, touchscreens, projectors, etc.) and one or more input devices (e.g., cameras (e.g., cameras facing downwards at the user's hands or forwards from the user's head (e.g., color sensors, infrared sensors, and other depth-sensing cameras)), controllers, touch-sensing surfaces, joysticks, buttons, etc.). In some embodiments, Method 10000 is executed by instructions stored in a non-temporary computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., control unit 110 in Figure 1A). Some operations of method 10000 are arbitrarily combined, and / or the order of some operations is arbitrarily changed.
[0216] In method 10000, the computer system displays a three-dimensional scene via a display generation component (10002), the three-dimensional scene comprising a first set of physical elements (e.g., physical objects or their representations shown in Figure 7G) (e.g., physical objects visible through the transparency of the display generation component, or physical objects represented by images of those objects in a camera view of the physical environment of the physical objects, where the position of each physical element in the three-dimensional scene corresponds to the position of each physical object in the physical environment surrounding the display generation component) and a first quantity of virtual elements (e.g., no virtual objects, or very simple virtual objects representing user interface elements and controls). The first set of physical elements includes at least physical elements corresponding to physical objects of a first class (e.g., walls or walls directly facing the display generation component, windows, etc.) and physical elements corresponding to physical objects of a second class (side walls distinct from walls directly facing the display generation component, ceilings and floors distinct from walls, walls distinct from windows, physical objects in a room, vertical physical surfaces inside a room, horizontal surfaces inside a room, surfaces larger than a preset threshold, surfaces of actual furniture in a room, etc.). While displaying a three-dimensional scene with a first quantity of virtual elements via the display generation component, the computer system detects a sequence of two or more user inputs (e.g., a sequence of two or more swipe inputs, a sequence of two or more snaps, inputs corresponding to the user putting on the HMD, then the user releasing the HMD from their hands, then the user sitting down with the HMD on their head) (10004) (e.g., the two or more user inputs are separate from user inputs dragging and / or dropping a particular virtual object into the three-dimensional scene (by input focus on that particular virtual object)).In response to detecting a sequence of two or more user inputs, the computer system continuously increases the amount of virtual elements displayed in the three-dimensional scene according to the sequence of two or more user inputs (10006) (for example, continuously increasing the immersion of the three-dimensional scene by replacing additional classes of physical elements in the three-dimensional scene in response to a sequence of user inputs of the same type or a sequence of related inputs). Specifically, in response to detecting a first user input in a sequence of two or more user inputs (e.g., an input by hand 7200 in Figure 7G), the computer system displays the three-dimensional scene with at least a first subset of one or more physical elements of a first set (e.g., some, but not all, of one or more physical elements of the first set are obscured or blocked by newly added virtual elements) and a second amount of virtual elements (e.g., virtual object 7402 in Figures 7H and 7I). The second amount of virtual elements occupies a larger portion of the three-dimensional scene than the first amount of virtual elements, including the first portion of the three-dimensional scene that was occupied by a first class of physical elements (e.g., walls) before the detection of the first user input (e.g., displaying virtual elements such as virtual landscapes or virtual windows that block the view of the first set of physical surfaces (e.g., walls) in the three-dimensional scene). In addition, in response to detecting a second user input in a sequence of two or more user inputs (e.g., an input by hand 7200 in Figure 7I), and in accordance with the determination that the second user input follows the first user input and satisfies the first criterion, the computer system displays a three-dimensional scene having at least a second subset of one or more physical elements of a first set (e.g., more or all of one or more physical elements of the first set are obscured or blocked by newly added virtual elements) and a third amount of virtual elements (e.g., virtual objects 7402 and 7406 in Figures 7J and 7K).The third quantity of virtual elements occupies a larger portion of the three-dimensional scene than the second quantity of virtual elements, including the first portion of the three-dimensional scene occupied by the first class of physical elements before the detection of the first user input, and the second portion of the three-dimensional scene occupied by the second class of physical elements before the detection of the second user input (for example, continuing to display virtual elements such as virtual scenery or virtual windows that obstruct the view of the first set of physical surfaces in the three-dimensional scene (e.g., walls), and displaying additional virtual elements such as virtual decorations or virtual surfaces that obstruct the view of the second set of physical surfaces (e.g., tabletops, shelves, or equipment surfaces)). This is illustrated, for example, in Figures 7G-7L.
[0217] In some embodiments, displaying a second quantity of virtual elements in response to detecting a first user input in a sequence of two or more user inputs includes displaying a first animation transition that gradually replaces a first class of physical elements whose quantity increases in the three-dimensional scene with virtual elements (new virtual elements and / or extensions of existing virtual elements) (e.g., replacing the display of an object that becomes visible via pass-through video, or obscuring an object that becomes directly visible through a transparent or partially transparent display). Displaying a third quantity of virtual elements in response to detecting a second user input in a sequence of two or more user inputs includes displaying a second animation transition that gradually replaces a second class of physical elements whose quantity increases in the three-dimensional scene with virtual elements (e.g., new virtual elements and / or extensions of existing virtual elements) while displaying a first class of physical elements in place of existing virtual elements in the three-dimensional scene (e.g., virtual elements of a second quantity). For example, in response to a first input (e.g., a first swipe input on the controller or the user's hand), the device replaces the view of a first physical wall visible in the 3D scene (e.g., a wall directly facing a display-generating component) with a virtual forest landscape, leaving other physical walls, physical ceilings, and physical floors visible in the 3D scene. When replacing the view of the first physical wall, the device displays an animated transition that gradually fades in with the virtual forest landscape. In response to a second input (e.g., a second swipe input on the controller or the user's hand), the device replaces the views of the remaining physical walls visible in the 3D scene (e.g., walls not directly facing a display-generating component) with a virtual forest landscape that extends from the already visible portion in the 3D scene, leaving only the physical ceiling and physical floors visible in the 3D scene. When replacing the views of the remaining physical walls, the device displays an animated transition that gradually extends the existing view of the virtual forest from the position of the first physical wall to the rest of the wall.In some embodiments, in response to a third input (e.g., a third swipe input on the controller or the user's hand), the device gradually replaces the view of the ceiling (and optionally, the floor), which is still visible in the three-dimensional scene, with a virtual landscape of the forest that gradually extends from the position of the surrounding physical walls toward the center of the ceiling (and optionally toward the center of the floor) from the existing view of the virtual forest (e.g., showing a portion of the virtual sky visible from a cut in the virtual forest) (e.g., showing a portion of the ground visible from a cut in the virtual forest). In response to a fourth input (e.g., a fourth swipe input on the controller or the user's hand), the device replaces the view of other physical objects, which are still visible in the three-dimensional scene, with a virtual overlay that gradually fades in onto the surface of the physical objects and becomes progressively opaque and saturated.
[0218] In some embodiments, when the amount of virtual elements is continuously increased according to a sequence of two or more user inputs, the computer system, in response to detecting a third user input in a sequence of two or more user inputs, displays the three-dimensional scene with a fourth amount of virtual elements, according to the determination that the third user input follows the second user input and satisfies a first criterion. The fourth amount of virtual elements occupies a larger portion of the three-dimensional scene than the third amount of virtual elements, including a first portion of the three-dimensional scene occupied by a first class of physical elements (e.g., a physical window or a wall facing a display-generating component) before the detection of the first user input, a second portion of the three-dimensional scene occupied by a second class of physical elements (e.g., a wall or a wall not facing a display-generating component) before the detection of the second user input, and a third portion of the three-dimensional scene occupied by a third class of physical elements (e.g., a physical object in a room) before the detection of the third user input (e.g., the fourth amount occupies the entire three-dimensional scene).
[0219] In some embodiments, in response to detecting a second user input in a sequence of two or more user inputs, and according to a determination that the second user input follows a first user input and satisfies a first criterion, the computer system displays a third animation transition between the display of a second quantity of virtual elements and the display of a third quantity of virtual elements. In some embodiments, the rendering of the second quantity of virtual elements is more artificial and less realistic, while the rendering of the third quantity of virtual elements (including the previously displayed second quantity of virtual elements and additional virtual elements) is more realistic and presents a more immersive computer-generated reality experience.
[0220] In some embodiments, the second quantity of virtual elements includes a view to a first virtual environment (e.g., a virtual window showing a scene at a different geographic location (e.g., a real-time video feed or a simulated scene)) which is represented by at least one or more physical elements of the first set. The view to the first virtual environment has a first set of values for first display characteristics (e.g., luminance distribution, color, hue, etc.) of the portion of the first virtual environment represented in the view (e.g., the virtual window shows pink morning light reflected from the top of a snow-covered mountain). The computer system modifies the visual appearance of at least some of the first subset of one or more physical elements of the first set according to a first set of values for a first display characteristic of the portion of the first virtual environment represented in a view of the first virtual environment (for example, the correspondence between the first set of values for the first display characteristic of the view of the first virtual environment shown in the virtual window and the changes in the visual appearance of the physical elements of the first subset is based on simulated physical principles such as the physical light propagation principle that determines how light is reflected and transmitted and how objects are illuminated by this light in the real world, the actual or simulated surface characteristics of the physical elements of the first subset, and the relative position of the virtual window to the physical elements of the first subset in the three-dimensional scene).
[0221] In some embodiments, while displaying a second amount of virtual elements (e.g., virtual windows showing scenes at different geographic locations (e.g., real-time video feeds or simulated scenes)) including a view to a first virtual environment displayed in a first subset of one or more physical elements of a first set, the computer system detects input that satisfies a second criterion (e.g., a criterion for displaying a navigation menu for changing the view to the virtual environment without changing the level of immersion) (e.g., a criterion for detecting a long-press gesture by the user's finger or hand). In response to detecting input that satisfies a second criterion separate from the first criterion (e.g., at least a time threshold, a sustained long-press input), the computer system displays a number of selectable options for changing the view to the first virtual environment (e.g., including menu options for changing the virtual environment represented in the virtual window (e.g., by changing location, time, lighting, weather conditions, zoom level, viewpoint, season, date, etc.)). In some embodiments, the computer system detects input indicating a selection of one of the displayed selectable options, and in response, the computer system replaces the view to a first virtual environment with a view to a second virtual environment different from the first virtual environment (e.g., an ocean or a cave), or updates the view to show the first virtual environment with at least one modified parameter that changes the appearance of the first virtual environment (e.g., time, season, date, location, zoom level, field of view, etc.).
[0222] In some embodiments, while displaying a second amount of virtual elements (e.g., virtual windows showing scenes at different geographic locations (e.g., real-time video feeds or simulated scenes)) including a view to a first virtual environment displayed in at least one or more physical elements of a first set, the computer system detects inputs that satisfy a third criterion (e.g., a criterion for changing the view to the virtual environment without changing the level of immersion) (e.g., a criterion for detecting a swipe gesture by the user's finger or hand). In response to detecting input that satisfies the third criterion, the computer system replaces the view to the first virtual environment with a view to a second virtual environment separate from the first virtual environment (e.g., an ocean or a cave). In some embodiments, as the content of the view changes (for example, along with changes in time, location, zoom level, field of view, season, etc.), the computer system also modifies the visual appearance of at least some of a portion of a first subset of one or more physical elements of a first set according to the modified values of a first display characteristic of the portion of the virtual environment represented in the content of the view (for example, the correspondence between a first set of values for a first display characteristic of the view of the virtual environment shown in the virtual window and the changes in the visual appearance of the physical elements of the first subset is based on simulated physical principles such as the physical light propagation principle that determines how light is reflected and transmitted and how objects are illuminated by this light in the real world, the actual or simulated surface characteristics of the physical elements of the first subset, and the relative position of the virtual window to the physical elements of the first subset in a three-dimensional scene).
[0223] In some embodiments, while displaying a second amount of virtual elements (e.g., virtual windows showing scenes at different geographic locations (e.g., real-time video feeds or simulated scenes)) including a view to a first virtual environment displayed in at least one or more physical elements of a first set, the computer system detects input that satisfies a third criterion (e.g., a criterion for changing the view to the virtual environment without changing the level of immersion) (e.g., a criterion for detecting a swipe gesture by the user's finger or hand). In response to detecting input that satisfies the third criterion, the computer system updates the view to show the first virtual environment with at least one modified parameter (e.g., time, season, date, location, zoom level, field of view, etc.) that modifies the appearance of the first virtual environment. In some embodiments, as the content of the view changes (for example, along with changes in time, location, zoom level, field of view, season, etc.), the computer system also modifies the visual appearance of at least some of a portion of a first subset of one or more physical elements of a first set according to the modified values of a first display characteristic of the portion of the virtual environment represented in the content of the view (for example, the correspondence between a first set of values for a first display characteristic of the view of the virtual environment shown in the virtual window and the changes in the visual appearance of the physical elements of the first subset is based on simulated physical principles such as the physical light propagation principle that determines how light is reflected and transmitted and how objects are illuminated by this light in the real world, the actual or simulated surface characteristics of the physical elements of the first subset, and the relative position of the virtual window to the physical elements of the first subset in a three-dimensional scene).
[0224] In some embodiments, the first criterion includes a first directional criterion (e.g., the input is a horizontal swipe input), and the second criterion includes a second directional criterion separate from the first directional criterion (e.g., the input is a vertical swipe input). For example, in some embodiments, a vertical swipe gesture increases or decreases immersion (e.g., increases or decreases the amount of virtual elements in a three-dimensional scene), while a horizontal swipe gesture changes the view represented within the virtual window without changing the size of the window, or changes the immersion (e.g., without changing the amount of virtual elements in a three-dimensional scene).
[0225] In some embodiments, displaying a first quantity of virtual elements includes displaying a first virtual window in a three-dimensional scene, displaying a second quantity of virtual elements includes extending the first virtual window in a three-dimensional scene, and displaying a third quantity of virtual elements includes replacing views of one or more physical walls with virtual elements. In some embodiments, additional user inputs in a sequence of two or more user inputs cause an additional quantity of virtual elements to be introduced into the three-dimensional scene, occupying portions of the scene previously occupied by physical elements. For example, a third input satisfying a first criterion replaces several remaining walls and ceilings with virtual elements. A fourth input satisfying a first criterion replaces the floor with a virtual element.
[0226] In some embodiments, a sequence of two or more user inputs includes repeated inputs of a first input type (e.g., the same input type, such as a vertical / upward swipe input).
[0227] In some embodiments, a sequence of two or more user inputs includes a continuous portion of a continuous input (e.g., a vertical / upward swipe input that includes a continuous movement in a predetermined direction starting from a first position and passing through a plurality of threshold positions / distances, or a press input that continuously increases in intensity beyond a plurality of intensity thresholds), and each portion of the continuous input corresponds to each user input of the sequence of two or more user inputs (e.g., by satisfying the corresponding input threshold of the plurality of input thresholds).
[0228] In some embodiments, a first subset of one or more physical elements of a first set includes at least the walls and floor of the physical environment, and a second subset of one or more physical elements of the first set includes the floor of the physical environment but does not include the walls of the physical environment. For example, in some embodiments, a virtual element replaces one or more walls of the physical environment represented in the three-dimensional scene, but does not replace the floor of the physical environment.
[0229] In some embodiments, a first subset of one or more physical elements of a first set includes at least walls and one or more pieces of furniture in the physical environment, while a second subset of one or more physical elements of the first set includes one or more pieces of furniture in the physical environment, but without including walls in the physical environment. For example, in some embodiments, a virtual element replaces one or more walls of the physical environment represented in the three-dimensional scene, but does not replace at least some of the furniture in the physical environment.
[0230] It should be understood that the specific order of operations described in Figure 10 is merely an example, and is not intended to indicate that the described order is the only order in which the operations can be performed. Those skilled in the art will recognize various methods for rearranging the operations described herein. In addition, it should be noted that details of other processes described herein with respect to other methods described herein (e.g., methods 8000, 9000, 11000, and 12000) are also applicable in a manner similar to method 10000 described above in relation to Figure 10. For example, the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to method 10000 optionally have one or more characteristics of the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to other methods described herein (e.g., methods 8000, 9000, 11000, and 12000). For brevity, those details will not be repeated here.
[0231] Figure 11 is a flowchart of Method 11000, which facilitates user interaction with a computer-generated environment (for example, by utilizing interaction with a physical surface to control a device or interact with a computer-generated environment) according to several embodiments. In some embodiments, Method 11000 is executed in a computer system (e.g., computer system 101 in Figure 1) that includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more input devices (e.g., a camera (e.g., a camera facing downwards at the user's hands or facing forward from the user's head (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras), a controller, a touch-sensing surface, a joystick, a button, etc.). In some embodiments, Method 11000 is stored in a non-temporary computer-readable storage medium and executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., a control unit 110 in Figure 1A). Some operations of Method 11000 are optionally combined, and / or the order of some operations is optionally changed.
[0232] In Method 11000, a computer system display displays a three-dimensional scene via a display generation component (11002), the three-dimensional scene including at least a first physical object (e.g., box 7502 or box 7504 in Figure 7M) or a representation thereof (e.g., representation 7502' or representation 7504' in Figure 7N). The first physical object has at least a first physical (substantially flat and / or smooth) surface (e.g., the first physical object is visible in the three-dimensional scene through a camera or transparent display). Each position of the first physical object or its representation in the three-dimensional scene corresponds to each position of the first physical object in the physical environment surrounding the display generation component (e.g., the first physical object is visible through a transparent pass-through portion of a head-up display or HMD, or the representation of the first physical object includes an image of the first physical object in a camera view of the physical environment displayed on the display or HMD). While displaying a three-dimensional scene, the computer system detects when a first interaction criterion is met (11004), the first interaction criterion includes a first criterion that is met when a first level of user interaction between the user and a first physical object is detected (for example, when the user's gaze is directed towards the first physical object without any other gesture or action indicating that the user wants to perform an action toward the first physical object, such as a hand movement or verbal command). In response to detecting that the first interaction criterion is met, the computer system displays a first user interface (e.g., the first user interface 7510 in Figure 7O or the first user interface 7516 in Figure 7P) (e.g., a simplified user interface or an information interface) (11006) (e.g., the first user interface is displayed on or superimposed on the first physical surface of the first physical object or at least a portion thereof).While a first user interface is displayed at a position corresponding to the position of a first physical surface or representation of a first physical object in a three-dimensional scene, the computer system detects that a second interaction criterion is met (11008), the second interaction criterion includes a second criterion that is met when a second level of user interaction is detected that is higher than a first level of user interaction between the user and the first physical object (for example, when the user or the user's hand approaches the first physical object while the user's gaze is still directed towards the first physical object) (for example, a level of user interaction that satisfies the second criterion also satisfies the first criterion, but a level of user interaction that satisfies the first criterion does not satisfy the second criterion). In response to detecting that the second interaction criterion is met, the computer system replaces the first user interface with a second user interface (e.g., the second user interface 7510' in Figure 7Q or the second user interface 7512' in Figure 7R) (e.g., an extended user interface or a user interface with control elements) at a location corresponding to the first physical surface or its representation of the first physical object in the three-dimensional scene (e.g., the second user interface corresponds to an extended user interface corresponding to the first physical object in comparison to the first physical object). In some embodiments, when the user's hand is detected near the first physical object (e.g., hover input is detected), the computer system displays more information (past / future songs, extended controls) in a third user interface that replaces the second user interface. In some embodiments, the first user interface includes a keyboard indication, the second user interface includes a keyboard with keys for text input, and the first physical surface is the tabletop of a physical table. The keyboard indication appears when the user looks at the tabletop, and the keyboard appears when the user looks at the tabletop and raises their hands above the tabletop in a typing position.In some embodiments, keyboard keys pop up from their positions in a three-dimensional scene corresponding to the tabletop when the user raises their hands above the tabletop. In some embodiments, when the user's fingers press or touch the tabletop, the keys are pressed down at positions corresponding to the touched locations on the tabletop and optionally appear to enlarge. When the user's fingers are lifted off the tabletop, the keys are restored to their original size.
[0233] In some embodiments, while displaying a second user interface at a location corresponding to the location of a first physical surface or representation of a first physical object in a three-dimensional scene, the computer system detects that a first interaction criterion has been met (for example, the level of user interaction returns to the first level of user interaction). In response to detecting that the first interaction criterion has been met after the display of the second user interface, the computer system replaces the display of the second user interface at the location corresponding to the location of the first physical surface or representation of the first physical object in the three-dimensional scene with the display of the first user interface. For example, the display of the extended user interface is stopped once the level of user interaction falls below the threshold level required to display the extended user interface. In some embodiments, if the level of user interaction falls further and the first interaction criterion is also not met, the computer system also stops the display of the first user interface.
[0234] In some embodiments, while displaying a first user interface (e.g., a media playback user interface) at a position corresponding to the location of a first physical surface or representation of a first physical object (e.g., a speaker) in a three-dimensional scene, the computer detects that a third interaction criterion has been met, the third interaction criterion includes a third criterion that is met when a first level of user interaction is detected between the user and a second physical object (e.g., a smart lamp) separate from the first physical object (e.g., when the user or the user's hand is not moving, but the user's line of sight moves from the first physical object to the second physical object). In response to detecting that the third interaction criterion has been met, the computer system stops displaying the first user interface (e.g., a media playback user interface) at a position corresponding to the location of the first physical surface or representation of the first physical object (e.g., a speaker) in the three-dimensional scene, and the computer system displays a third user interface (e.g., a lighting control user interface) at a position corresponding to the location of a second physical surface or representation of a second physical object (e.g., a smart lamp) in the three-dimensional scene. For example, when the user's gaze shifts from a first physical object to a second physical object, while the user's hand is suspended in the air without moving near either the first or second physical object, the computer system stops displaying the user interface corresponding to the first physical object that overlaps the surface of the first physical object, and instead displays the user interface corresponding to the second physical object that overlaps the surface of the second physical object in the three-dimensional scene.
[0235] In some embodiments, while a first user interface is displayed at a position corresponding to a first physical surface or representation of a first physical object in a three-dimensional scene, the computer system detects a first input that satisfies a first action criterion, the first action criterion corresponding to the activation of a first option included in the first user interface (for example, the first activation criterion is a criterion for detecting a tap input). In response to detecting a first input that satisfies the first action criterion while the first user interface is displayed, the computer system performs a first action corresponding to a first option included in the first user interface (for example, activating the play / pause function of a media player associated with the first physical object (e.g., a speaker or stereo)). In some embodiments, while a first user interface is displayed at a position corresponding to a first physical surface or representation of a first physical object in a three-dimensional scene, the computer system detects a second input that satisfies a second action criterion, the second action criterion corresponding to the activation of a second option included in the first user interface (for example, the second action criterion is a criterion for detecting a swipe input or a criterion for detecting a twist input), and in response to detecting a second input that satisfies the second action criterion while the first user interface is displayed, the computer system performs a second action corresponding to the second option included in the first user interface (for example, activating the fast-forward or rewind function of a media player associated with the first physical object (e.g., a speaker or stereo), or adjusting the volume or output level of the first physical object).
[0236] In some embodiments, while a second user interface is displayed at a position corresponding to a first physical surface or representation of a first physical object in a three-dimensional scene, the computer system detects a second input that satisfies a third action criterion, the third action criterion corresponding to the activation of a third option included in the second user interface (for example, the third action criterion is a criterion for detecting a tap input along with gaze input directed at a first user interface object included in the second user interface). In response to detecting a first input that satisfies the third action criterion while the second user interface is displayed, the computer system performs a third action corresponding to a third option included in the second user interface (for example, switching to a different album of a media player associated with a first physical object (e.g., a speaker or stereo)). In some embodiments, while a second user interface is displayed at a position corresponding to the position of a first physical surface or representation of a first physical object in a three-dimensional scene, the computer system detects a fourth input that satisfies a fourth action criterion, the fourth action criterion corresponding to the activation of a fourth option included in the second user interface (for example, the fourth action criterion is a criterion for detecting a swipe input accompanied by gaze input directed at a second user interface object included in the second user interface), and in response to detecting a fourth input that satisfies the fourth action criterion while the second user interface is displayed, the computer system performs a fourth action corresponding to the fourth option included in the second user interface (for example, activating one or more other associated physical objects of the first physical object (for example, activating one or more other associated speakers), or sending the output of the first physical object to another physical object).
[0237] In some embodiments, the first physical object is a speaker, and the first user interface provides a first set of one or more playback control functions associated with the speaker (e.g., a play / pause control function, a fast-forward function, a rewind function, a stop function, etc.). In some embodiments, the first user interface includes user interface objects corresponding to these control functions. In some embodiments, the first user interface does not include user interface objects corresponding to at least some of the control functions provided to the first user interface at a given time, and the user interface objects displayed on the first user interface are selected in response to user input detected while the first user interface is displayed. For example, when a user provides a swipe input while the first user interface is displayed, the first user interface displays a fast-forward or rewind symbol depending on the direction of the swipe input. When a user provides a tap input while the first user interface is displayed, the first user interface displays a play / pause indicator depending on the current state of playback. When the user provides pinch and twist input with their fingers, the first user interface displays a volume control that adjusts the speaker volume level according to the direction of the twist input. In some embodiments, the first user interface also provides information that the user can select, such as a list of recently played or next songs / albums.
[0238] In some embodiments, the first user interface includes one or more notifications corresponding to a first physical object. For example, when a user has a first level of interaction with the first physical object (e.g., the user sees a speaker or smart lamp), the computer system displays one or more notifications that overlap the first physical surface of the first physical object (e.g., notifications related to the status or warnings corresponding to the speaker or smart lamp (e.g., "Low Battery," "Set Timer to 20 Minutes")).
[0239] In some embodiments, the second user interface includes a keyboard with multiple character keys for text input. For example, when the user has a second level of interaction with a first physical object (e.g., the user looks at the speaker and raises both hands), the computer displays a search interface, along with the keyboard, for the user to enter search keywords to search a music database associated with the speaker.
[0240] In some embodiments, the first user interface displays an indication of the internal state of the first physical object. For example, when the user has a first level of interaction with the first physical object (e.g., the user looks at a speaker or smart lamp), the computer system displays the internal state of the first physical object overlaid on the first physical surface of the first physical object (e.g., the name of the album / song currently playing, "low battery", "timer set to 20 minutes", etc.).
[0241] In some embodiments, the second user interface provides at least a subset of the functions or information provided to the first user interface, and includes at least one function or item of information not available in the first user interface. For example, when the user has a first level of interaction with the first physical object (e.g., the user looks at the speaker or smart lamp), the computer system displays the internal state of the first physical object overlaid on the first physical surface of the first physical object (e.g., the name of the album / song currently playing, "low battery", "timer set to 20 minutes", etc.), and when the user has a second level of interaction with the first physical object (e.g., the user looks at the speaker or smart lamp and makes a ready gesture to offer input by raising their hand, or approaches the first physical object), the computer system displays a user interface that displays the internal state of the first physical object, as well as one or more controls for changing the internal state of the first physical object (e.g., a control for changing the song / album currently playing, a control for routing the output to the associated speaker, etc.).
[0242] In some embodiments, while a first user interface is displayed at a position corresponding to the position of a first physical surface of a first physical object in a three-dimensional scene (for example, the first user interface is displayed on or overlapping at least a portion of the first physical surface or representation thereof of the first physical object), the computer system detects a user input that satisfies a fifth criterion (for example, a criterion for detecting a swipe input while eye-tracking input is focused on the first user interface) that responds to a request to remove the first user interface. In response to detecting a user input that satisfies the fifth criterion, the computer system stops displaying the first user interface (for example, without replacing the first user interface with a second user interface). Similarly, in some embodiments, while a second user interface is displayed at a position corresponding to the position of a first physical surface of a first physical object in a three-dimensional scene (for example, the second user interface is displayed on or overlapping at least a portion of the first physical surface or representation thereof of the first physical object), the computer system detects a user input that satisfies a sixth criterion (for example, a criterion for detecting a swipe input while eye-tracking input is focused on the second user interface) that responds to a request to remove the second user interface, and in response to detecting a user input that satisfies the sixth criterion, the computer system stops displaying the second user interface (for example, without replacing the second user interface with the first user interface).
[0243] In some embodiments, while displaying a first user interface or a second user interface at a position corresponding to the position of a first physical surface of a first physical object in a three-dimensional scene (the first / second user interface is displayed on or overlapping at least a portion of the first physical surface of the first physical object or its representation), the computer system detects user input on the first physical surface of the first physical object (e.g., using one or more sensors on the physical surface, such as a touch sensor or proximity sensor, and / or one or more sensors on a device, such as a camera or depth sensor). In response to detecting user input on the first physical surface of the first physical object, the computer system performs a first action corresponding to the first physical object, according to a determination that the user input on the first physical surface of the first physical object satisfies a sixth criterion (e.g., a first criterion set in each criterion set for detecting swipe input, tap input, long press input, or double tap input, etc.). Upon determining that user input on the first physical surface of the first physical object satisfies a sixth criterion (for example, a second set of criteria in a set of criteria for detecting swipe input, tap input, long press input, or double tap input, etc.), the computer system performs a second action corresponding to the first physical object, separate from the first action.
[0244] In some embodiments, while displaying a first or second user interface at a position corresponding to the position of a first physical surface of a first physical object in a three-dimensional scene (for example, the first / second user interface is displayed on or overlapping at least a portion of the first physical surface or representation thereof of the first physical object), the computer system detects a gesture input (e.g., a hand gesture in the air, on a controller, or on the user's hand) while the gaze input is directed to the first physical surface of the first physical object. In response to the detection of the gesture input while the gaze input is directed to the first physical surface of the first physical object, the computer system performs a third action (e.g., a function associated with a button) corresponding to the first physical object, according to the determination that the gesture input and the gaze input satisfy a seventh criterion (e.g., the gesture is a tap input while the gaze input is on a button in the user interface). Upon determining that the gesture input and gaze input satisfy the eighth criterion (for example, the gesture is a swipe input while the gaze input is on a slider in the user interface), the computer system performs a fourth action corresponding to a first physical object separate from the third action (for example, adjusting the value associated with the slider).
[0245] In some embodiments, while a first user interface or a second user interface is displayed at a position corresponding to the position of a first physical surface of a first physical object in a three-dimensional scene (for example, the first / second user interface is displayed on or superimposed on the first physical surface of the first physical object or at least a portion thereof), the computer system detects gesture input on a second physical surface of a second physical object separate from the first physical object (for example, the second physical object is a tabletop or controller near the user's hand) (using one or more sensors on the physical surface, such as a touch sensor or proximity sensor, and / or one or more sensors on a device, such as a camera or depth sensor) while the gaze input is directed to the first physical surface of the first physical object (for example, the first physical object is far away from the user's hand). In response to detecting a gesture input on the second physical surface of a second physical object while the gaze input is directed to the first physical surface of a first physical object, the computer system performs a fifth action corresponding to the first physical object (e.g., a function associated with a button) according to the determination that the gesture input and gaze input satisfy the ninth criterion (e.g., the gesture is a tap input while the gaze input is on a button in the user interface), and the computer system performs a sixth action corresponding to the first physical object, separate from the fifth action (e.g., adjusting a value associated with a slider) according to the determination that the gesture input and gaze input satisfy the tenth criterion (e.g., the gesture is a swipe input while the gaze input is on a slider in the user interface).
[0246] It should be understood that the specific order of operations described in Figure 11 is merely an example, and is not intended to indicate that the described order is the only order in which the operations can be performed. Those skilled in the art will recognize various methods for rearranging the operations described herein. In addition, it should be noted that details of other processes described herein with respect to other methods described herein (e.g., methods 8000, 9000, 10000, and 12000) are also applicable in a manner similar to method 11000 described above in relation to Figure 11. For example, the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to method 11000 optionally have one or more characteristics of the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to other methods described herein (e.g., methods 8000, 9000, 10000, and 12000). For brevity, those details will not be repeated here.
[0247] Figure 12 is a flowchart of Method 12000, which generates a computer-generated three-dimensional environment (including, for example, simulating visual interactions between a voice-based virtual assistant and physical and virtual objects within the environment) according to several embodiments. In some embodiments, Method 12000 is executed in a computer system (e.g., computer system 101 in Figure 1) that includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more input devices (e.g., a camera (e.g., a camera (e.g., a camera facing downwards at the user's hands or facing forward from the user's head (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras)), a controller, a touch-sensing surface, a joystick, a button, etc.). In some embodiments, Method 12000 is stored in a non-temporary computer-readable storage medium and executed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., a control unit 110 in Figure 1A). Some operations of Method 12000 are optionally combined, and / or the order of some operations is optionally changed.
[0248] In method 12000, the computer system displays a three-dimensional scene via a display generation component (12002), the three-dimensional scene includes a first physical object (e.g., the front surface 7312 of the furniture 7310 in Figure 7T) (e.g., the first physical object is visible in the three-dimensional scene via a camera or transparent display, (e.g., first visible and possessing intrinsic optical properties such as color, texture, reflectivity, and transparency)) and a first virtual object (e.g., the virtual object 7404 in Figure 7T) having at least a first virtual surface (e.g., a computer-rendered three-dimensional object having a computer-generated surface with simulated surface optical properties (e.g., simulated reflectivity, simulated surface texture, etc.), such as a computer-rendered three-dimensional vase or tabletop). While displaying a three-dimensional scene, the computer system detects a request to activate a voice-based virtual assistant, for example, as shown in Figure 7T (12004). In response to detecting the request to activate a voice-based virtual assistant (12006), the computer system activates a voice-based virtual assistant configured to receive voice commands (for example, to interact with the three-dimensional scene). The computer system also displays a visual representation of the voice-based virtual assistant in the three-dimensional scene (e.g., the luminescent ellipse 7602 in Figures 7U and 7W), which includes displaying a visual representation of the voice-based virtual assistant with a first set of values of a first display characteristic of the visual representation (e.g., color, or luminance) (e.g., single values, continuous values, or distinct and discrete values of different parts of the visual representation) (e.g., the luminescent ellipse 7602 has a first range of luminance and a first color).The computer system modifies the visual appearance of at least a portion of the first physical surface of a first physical object (e.g., the front 7312 of furniture 7310 in Figures 7U and 7W or its representation) and at least a portion of the first virtual surface of a first virtual object (e.g., the top surface of virtual object 7404 in Figures 7U and 7W) according to a first set of values of the first display characteristics of the visual representation of the voice-based virtual assistant (for example, the correspondence between the first set of values of the first display characteristics of the visual representation of the voice-based virtual assistant and the changes in the visual appearance of the first physical surface and the first virtual surface is based on simulated physical principles such as the principle of light propagation which determines how light is reflected and transmitted and how objects are illuminated by this light in the real world, the actual or simulated surface characteristics of the first physical surface and the first virtual surface, and the relative position of the virtual assistant with respect to the first physical surface and the first virtual surface). For example, as shown in Figure 7U, when the voice-based assistant's representation begins to emit light at a first level of brightness, the appearance of the front of the furniture 7310 is modified to appear as if it is illuminated by the simulated lighting emitted from the voice-based assistant's luminous representation. The simulated lighting on the front of the rectangular box is stronger / brighter closer to the voice-based assistant's luminous representation and weaker / dimmer further away from it. In some embodiments, the simulated lighting is generated according to the physical properties of the front of the rectangular box in the real world (e.g., surface texture, reflectivity, etc.) and the simulated distance between the voice-based assistant's luminous representation and the rectangular box in the three-dimensional scene. In some embodiments, the device also modifies the appearance of the three-dimensional scene by adding a simulated shadow next to the rectangular box (e.g., onto the physical wall behind the rectangular box) that is cast by the rectangular box under the simulated lighting by the voice-based virtual assistant's luminous representation. In some embodiments, in addition to modifying the appearance of the front of a rectangular box (for example, by using a translucent overlay at a position corresponding to a physical surface, or by directly modifying the displayed pixel values of a representation of a physical surface), the device also modifies the appearance of the top surface of a virtual elliptical object so that it appears to be illuminated by simulated lighting emitted from a luminous representation of a voice-based assistant.The simulated lighting on the top surface of the virtual ellipse object is stronger / brighter closer to the voice-based assistant's luminous representation and weaker / dimmer further away. In some embodiments, the simulated lighting is generated according to the simulated physical properties of the top surface of the virtual ellipse object (e.g., surface texture, reflectivity, etc.) and the simulated distance between the voice-based assistant's luminous representation and the virtual ellipse object in the three-dimensional scene. In some embodiments, the device also modifies the appearance of the three-dimensional scene by adding a simulated shadow next to the virtual ellipse object or by modifying existing simulated shadows cast by the virtual ellipse object according to the simulated lighting by the voice-based assistant's luminous representation.
[0249] In some embodiments, modifying the visual appearance of at least a portion of the first virtual faces of the first virtual object (e.g., the top surface of the virtual object 7404 in Figures 7U and 7W) according to a first set of values for the first display characteristics of the visual representation of the voice-based virtual assistant includes increasing the luminance of each of the at least portions of the first virtual faces of the first virtual object according to the increased luminance value of the visual representation of the voice-based virtual assistant (e.g., according to the increased luminance value corresponding to the portion of the visual representation of the voice-based virtual assistant facing a portion of the first virtual face of the first virtual object (e.g., a portion on a display generating component that may be invisible to the user)).
[0250] In some embodiments, modifying the visual appearance of at least a portion of the first virtual faces of the first virtual object (e.g., the top surface of the virtual object 7404 in Figures 7U and 7W) according to a first set of values for the first display characteristics of the visual representation of the voice-based virtual assistant includes changing the color of each of the at least portions of the first virtual faces of the first virtual object according to the changed color values of the visual representation of the voice-based virtual assistant (e.g., according to the changed color values corresponding to the portion of the visual representation of the voice-based virtual assistant facing a portion of the first virtual face of the first virtual object (e.g., the portion invisible to the user on the display generation component)).
[0251] In some embodiments, modifying the visual appearance of at least a portion of the first physical surface of a first physical object (e.g., the front 7312 of the furniture 7310 in Figures 7U and 7W, or its representation) according to a first set of values of the first display characteristics of the visual representation of a voice-based virtual assistant includes increasing the luminance of each portion of the three-dimensional scene corresponding to at least a portion of the first physical surface of the first physical object according to the increased luminance value of the visual representation of the voice-based virtual assistant (e.g., according to the increase in the luminance value corresponding to the portion of the visual representation of the voice-based virtual assistant facing a portion of the first physical surface of the first physical object (e.g., a portion that may be invisible to the user on a display generation component)).
[0252] In some embodiments, modifying the visual appearance of at least a portion of the first physical surface of a first physical object (e.g., the front 7312 of the furniture 7310 in Figures 7U and 7W, or its representation) according to a first set of values of the first display characteristics of the visual representation of a voice-based virtual assistant includes changing the color of each portion of the three-dimensional scene corresponding to at least a portion of the first physical surface of the first physical object according to the changed color values of the visual representation of the voice-based virtual assistant (e.g., according to the changed color values corresponding to the portion of the visual representation of the voice-based virtual assistant facing a portion of the first physical surface of the first physical object (e.g., a portion that may be invisible to the user on a display generation component)).
[0253] In some embodiments, in response to detecting a request to activate a voice-based virtual assistant, the computer system modifies the visual appearance of a peripheral area of a portion of the three-dimensional scene currently displayed via a display generation component according to a first set of values for a first display characteristic of the visual representation of the voice-based virtual assistant (e.g., increasing the brightness of the peripheral area, or changing its color or hue). For example, if the virtual assistant is represented by an luminous purple ellipse in the three-dimensional scene, the peripheral area of the user's field of view is displayed with an indistinct luminous border to indicate that a voice command to the voice-based virtual assistant is being executed on one or more objects in the portion of the three-dimensional scene currently within the user's field of view. For example, if the user looks around a room, the central area of the user's field of view is transparent and surrounded by a purple vignette, and the objects in the central area of the user's field of view are the targets of the voice command or provide context for the voice command detected by the voice-based virtual assistant (e.g., "Turn this on" or "Change this picture").
[0254] In some embodiments, detecting a request to activate a voice-based virtual assistant includes detecting eye-tracking input that satisfies a first criterion, the first criterion being satisfied when the eye-tracking input is directed to a position corresponding to a visual representation of the voice-based virtual assistant in a three-dimensional scene (for example, the virtual assistant is activated when the user gazes at a visual representation of the virtual assistant). In some embodiments, the first criterion also includes being satisfied when the eye-tracking input satisfies predetermined gaze fixation and duration thresholds. In some embodiments, the request to activate a voice-based virtual assistant includes a predetermined trigger command, "Hey, Assistant."
[0255] In some embodiments, displaying a visual representation of the voice-based virtual assistant (e.g., the luminous ellipse 7602 in Figures 7U and 7W) in a three-dimensional scene in response to detecting a request to activate the voice-based virtual assistant (e.g., in response to detecting eye-tracking input that satisfies a first criterion) includes moving the visual representation of the voice-based virtual assistant from a first position to a second position in the three-dimensional scene (e.g., when the user stares at a dormant virtual assistant, the virtual assistant pops up from its original position (e.g., in the center of the user's field of view, or slightly away from its original position to indicate that it has been activated)).
[0256] In some embodiments, displaying the voice-based virtual assistant in a three-dimensional scene (e.g., the luminous ellipse 7602 in Figures 7U and 7W) in response to detecting a request to activate the visual representation of the voice-based virtual assistant (e.g., in response to detecting eye-tracking input that satisfies a first criterion) includes resizing the visual representation of the voice-based virtual assistant in the three-dimensional scene (e.g., when the user stares at a dormant virtual assistant, the virtual assistant enlarges in size and then returns to its original size, or remains enlarged until it is deactivated again).
[0257] In some embodiments, displaying a visual representation of the voice-based virtual assistant in a three-dimensional scene (e.g., the luminous ellipse 7602 in Figures 7U and 7W) in response to detecting a request to activate the voice-based virtual assistant (e.g., in response to detecting eye-tracking input that satisfies a first criterion) includes changing a first set of values for a first display characteristic of the visual representation of the voice-based virtual assistant in the three-dimensional scene (e.g., when the user stares at a dormant virtual assistant, the virtual assistant emits light and / or has a different color or hue).
[0258] In some embodiments, in response to detecting a request to activate a voice-based virtual assistant (for example, in response to detecting eye-tracking input that satisfies a first criterion), the computer system modifies a second set of values for a first set of display characteristics of a portion of a three-dimensional scene at a location surrounding the visual representation of the voice-based virtual assistant (for example, obscuring (blurring, darkening, etc.) the background (for example, an area around the virtual assistant or an area around the entire screen) when the virtual assistant is invoked).
[0259] In some embodiments, detecting a request to activate a voice-based virtual assistant includes detecting eye-tracking input that satisfies a first criterion and voice input that satisfies a second criterion, the first criterion being satisfied when the eye-tracking input is directed to a location corresponding to a visual representation of the voice-based virtual assistant in a three-dimensional scene, and the second criterion being satisfied when voice input is detected while the eye-tracking input satisfies the first criterion (for example, the virtual assistant is activated when the user looks at a visual representation of the virtual assistant and speaks a voice command). In some embodiments, after the voice-based virtual assistant is activated, the device processes the voice input to determine a user command for the voice assistant and provides the user command to the virtual assistant as input to trigger the performance of the corresponding action by the virtual assistant. In some embodiments, if the eye-tracking input does not satisfy the first criterion or the voice input does not satisfy the second criterion, the virtual assistant does not perform an action corresponding to the voice command in the voice input.
[0260] In some embodiments, while displaying a visual representation of a voice-based virtual assistant in a three-dimensional scene (e.g., the luminous ellipse 7602 in Figures 7K and 7L), the computer system detects a first input corresponding to a request regarding the voice-based assistant and performs a first action (e.g., changing a photograph in a virtual photo frame in the scene, starting a communication session, starting an application, etc.), the first input lasting for a first period of time (e.g., the first input is a speech input, gaze input, gesture input, or a combination of two or more of the above). In response to detecting the first input, the computer system changes the first display characteristics of the visual representation of the voice-based virtual assistant during the first input from a first set of values (e.g., a single value, a continuous range of values for different parts of the visual representation, or discrete and distinct values for different parts of the visual representation) to a second set of values separate from the first set of values. In some embodiments, the device also modifies the visual appearance of at least a portion of the first physical surface of a first physical object (e.g., the front or representation thereof of furniture 7310 in Figures 7U and 7W) and at least a portion of the first virtual surface of a first virtual object (e.g., the top surface of virtual object 7602 in Figures 7U and 7W) according to a second set of values of the first display characteristics of the visual representation of the voice-based virtual assistant, as the values of the first display characteristics of the visual representation of the voice-based virtual assistant change during a first input. For example, while the user is speaking to the virtual assistant, the visual representation of the virtual assistant may emit light in pulsating light, various colors, or dynamic color / light patterns.
[0261] In some embodiments, while displaying a visual representation of a voice-based virtual assistant in a three-dimensional scene (e.g., the luminous ellipse 7602 in Figures 7U and 7W), the computer system detects a second input corresponding to a request regarding the voice-based assistant and performs a second action (e.g., changing a photograph in a virtual photographic frame in the scene, starting a communication session, starting an application, etc.) (for example, the second input is a speech input, gaze input, gesture input, or a combination of two or more of the above). In response to detecting the second input, the computer system begins to perform the second action (e.g., launching an application, playing a media file, generating an audio output such as a request for additional information or an answer to a question). The computer system also, while performing the second action, changes the first display characteristics of the visual representation of the voice-based virtual assistant from a first set of values (e.g., a single value, a continuous range of values for different parts of the visual representation, or distinct and discrete values for different parts of the visual representation) to a third set of values distinct from the first set of values. In some embodiments, the device also modifies the visual appearance of at least a portion of the first physical surface of a first physical object (e.g., the front 7312 of furniture 7310 in Figures 7U and 7W, or its representation) and at least a portion of the first virtual surface of a first virtual object (e.g., the top surface of virtual object 7404 in Figures 7U and 7W) according to a third set of values of the first display characteristics of the visual representation of the voice-based virtual assistant, as the values of the first display characteristics of the visual representation of the voice-based virtual assistant change during the performance of a second action by the virtual assistant. For example, while the user is speaking to the virtual assistant, the vis...
Claims
1. It is a method, In a computer system comprising one or more display generation components and one or more input devices, Displaying a three-dimensional scene via one or more display generation components, wherein the three-dimensional scene includes a first set of physical elements and a first quantity of virtual elements. While displaying the three-dimensional scene having a first amount of virtual elements via one or more display generation components, the detection of a sequence of two or more user inputs, including a first user input and a second user input following the first user input, wherein the first user input and the second user input satisfy a first criterion, and the user input satisfying the first criterion corresponds to a request to increase the current level of immersion in which the three-dimensional scene is displayed. In response to detecting a sequence of two or more user inputs, including detecting the first user input followed by the second user input, the amount of virtual elements displayed in the three-dimensional scene is continuously increased according to the sequence of two or more user inputs. In response to detecting the first user input in the sequence of two or more user inputs, a first animation transition is displayed in which an increasing portion of a first region of the three-dimensional scene occupied by one or more physical elements of the first set is gradually replaced with virtual elements; and at the end of the first animation transition, a first subset of at least one or more physical elements of the first set and a second amount of virtual elements are displayed in the three-dimensional scene, wherein the second amount of virtual elements occupies a larger portion of the three-dimensional scene than the first amount of virtual elements. In response to detecting the second user input in the sequence of two or more user inputs, a second animation transition is displayed which gradually replaces an increasing portion of a second area of the three-dimensional scene occupied by at least one or more physical elements of the first set with virtual elements, and at the end of the second animation transition, at least the second subset of one or more physical elements of the first set and a third amount of virtual elements are displayed in the three-dimensional scene, the third amount of virtual elements occupying a larger portion of the three-dimensional scene than the second amount of virtual elements, including displaying an increasing portion. Displaying the first animation transition includes gradually expanding a third region of the three-dimensional scene occupied by virtual elements to replace an increasing portion of the first region of the three-dimensional scene occupied by one or more physical elements of the first set, A method for displaying the second animation transition, comprising gradually expanding a fourth region of the three-dimensional scene occupied by virtual elements to replace an increasing portion of the second region of the three-dimensional scene occupied by the first subset of one or more physical elements of the first set.
2. The method according to claim 1, Detecting the consecutive user inputs of the sequence of two or more user inputs includes detecting a third user input that follows the second user input, wherein the third user input satisfies the first criterion. A method for continuously increasing the amount of virtual elements displayed in the three-dimensional scene in accordance with the sequential user inputs of the sequence of two or more user inputs, which includes displaying a fourth amount of virtual elements that occupies a larger portion of the three-dimensional scene than the third amount of virtual elements in response to the detection of a third user input in the sequence of two or more user inputs.
3. A method according to any one of claims 1 to 2, Displaying the first animation transition includes increasing the opacity of a first subset of the first amount of virtual elements, A method for displaying the second animation transition, comprising increasing the opacity of a second subset of the second amount of virtual elements.
4. A method, In a computer system comprising one or more display generation components and one or more input devices, Displaying a three-dimensional scene via one or more display generation components, wherein the three-dimensional scene includes a first set of physical elements and a first quantity of virtual elements. While displaying the three-dimensional scene having a first amount of virtual elements via one or more display generation components, the detection of a sequence of two or more user inputs, including a first user input and a second user input following the first user input, wherein the first user input and the second user input satisfy a first criterion, and the user input satisfying the first criterion corresponds to a request to increase the current level of immersion in which the three-dimensional scene is displayed. In response to detecting a sequence of two or more user inputs, including detecting the first user input followed by the second user input, the amount of virtual elements displayed in the three-dimensional scene is continuously increased according to the sequence of two or more user inputs. In response to detecting the first user input in the sequence of two or more user inputs, a first animation transition is displayed in which an increasing portion of a first region of the three-dimensional scene occupied by one or more physical elements of the first set is gradually replaced with virtual elements; and at the end of the first animation transition, a first subset of at least one or more physical elements of the first set and a second amount of virtual elements are displayed in the three-dimensional scene, wherein the second amount of virtual elements occupies a larger portion of the three-dimensional scene than the first amount of virtual elements. In response to detecting the second user input in the sequence of two or more user inputs, a second animation transition is displayed which gradually replaces an increasing portion of a second area of the three-dimensional scene occupied by at least one or more physical elements of the first set with virtual elements, and at the end of the second animation transition, at least the second subset of one or more physical elements of the first set and a third amount of virtual elements are displayed in the three-dimensional scene, the third amount of virtual elements occupying a larger portion of the three-dimensional scene than the second amount of virtual elements, including displaying an increasing portion. The second quantity of virtual elements includes a view to a first virtual environment displayed together with the first subset of one or more physical elements of the first set, the view to the first virtual environment having a first set of values for a first display characteristic of a portion of the first virtual environment displayed in the view. The method comprises modifying the visual appearance of at least a portion of a first subset of one or more physical elements of a first set according to a first set of values for the first display characteristics of a portion of the first virtual environment as displayed in the view to the first virtual environment.
5. The method according to claim 4, While displaying the three-dimensional scene which includes the first subset of one or more physical elements of the first set and the second amount of virtual elements, it is possible to detect that the content of the view to the first virtual environment has changed, A method comprising detecting that the content of the view to the first virtual environment has changed, and modifying the visual appearance of at least a portion of a first subset of one or more physical elements of the first set in accordance with a change in a first set of values for the first display characteristics of the portion of the first virtual environment displayed in the view to the first virtual environment.
6. The method according to claim 4, Detecting user input that satisfies a second criterion different from the first criterion while displaying a second amount of virtual elements, including the view to the first virtual environment displayed together with at least one or more physical elements of the first set, A method comprising: detecting user input that satisfies the second criterion, displaying a number of selectable options for changing the view to the first virtual environment.
7. The method according to claim 6, While displaying the multiple selectable options for changing the view to the first virtual environment, the system detects user input to select a first selectable option from the multiple selectable options, A method comprising: detecting the user input of selecting the first selectable option, replacing the view to the first virtual environment with a view to a second virtual environment different from the first virtual environment.
8. The method according to claim 7, A method comprising maintaining the current level of immersion to which the three-dimensional scene is displayed when replacing the view to the first virtual environment with the view to the second virtual environment in response to detecting the user input of selecting the first selectable option.
9. A method according to any one of claims 1 to 8, wherein the sequence of two or more user inputs includes repeated input of a first input type.
10. A method, In a computer system comprising one or more display generation components and one or more input devices, Displaying a three-dimensional scene via one or more display generation components, wherein the three-dimensional scene includes a first set of physical elements and a first quantity of virtual elements. While displaying the three-dimensional scene having a first amount of virtual elements via one or more display generation components, the detection of a sequence of two or more user inputs, including a first user input and a second user input following the first user input, wherein the first user input and the second user input satisfy a first criterion, and the user input satisfying the first criterion corresponds to a request to increase the current level of immersion in which the three-dimensional scene is displayed. In response to detecting a sequence of two or more user inputs, including detecting the first user input followed by the second user input, the amount of virtual elements displayed in the three-dimensional scene is continuously increased according to the sequence of two or more user inputs. In response to detecting the first user input in the sequence of two or more user inputs, a first animation transition is displayed in which an increasing portion of a first region of the three-dimensional scene occupied by one or more physical elements of the first set is gradually replaced with virtual elements; and at the end of the first animation transition, a first subset of at least one or more physical elements of the first set and a second amount of virtual elements are displayed in the three-dimensional scene, wherein the second amount of virtual elements occupies a larger portion of the three-dimensional scene than the first amount of virtual elements. In response to detecting the second user input in the sequence of two or more user inputs, a second animation transition is displayed which gradually replaces an increasing portion of a second area of the three-dimensional scene occupied by at least one or more physical elements of the first set with virtual elements, and at the end of the second animation transition, at least the second subset of one or more physical elements of the first set and a third amount of virtual elements are displayed in the three-dimensional scene, the third amount of virtual elements occupying a larger portion of the three-dimensional scene than the second amount of virtual elements, including displaying an increasing portion. A method wherein the sequence of two or more user inputs includes a sequence of inputs having a first portion of the sequence of inputs corresponding to the first user input of the sequence of two or more user inputs, and a second portion of the sequence of inputs corresponding to the second user input of the sequence of two or more user inputs.
11. A computer program that, when executed by a computer system having one or more processors and a display generation component, causes the computer system to perform the method described in any one of claims 1 to 10.
12. A computer system, One or more processors, Display generation component, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Information processing apparatus, space information use system, information processing method, control program, and recording medium
JP2013250897A
Systems and methods for viewport-based augmented reality haptic effects
JP2015212946A
Virtual reality display system, virtual reality display method, and computer program
JP2017004457A
Mediated reality
WO2017009529A1