Method for Providing an Immersive Experience in an Environment
The computer system with advanced input devices and interfaces addresses inefficiencies in augmented and virtual reality interactions by reducing input complexity and providing intuitive feedback, enhancing user experience and energy efficiency.
Patent Information
- Application Number
- JP2023562859
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-11
- Filing Date
- 2022-04-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-04-13
AI Technical Summary
Existing methods for interacting with augmented and virtual reality environments are inefficient, complex, and place a significant cognitive burden on users, often requiring multiple inputs and providing insufficient feedback, leading to wasted energy and degraded user experience.
A computer system with enhanced input devices and interfaces, including touch-sensitive displays, eye-tracking, hand-tracking, and haptic output generators, that reduce the number and complexity of user inputs by providing intuitive feedback and adapting the virtual environment to user movements.
The system enhances user interaction efficiency, reduces errors, and conserves energy by minimizing inputs and improving feedback, thereby improving the overall user experience in augmented and virtual reality environments.
Smart Images

Figure 0007713533000001 
Figure 0007713533000002 
Figure 0007713533000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 174,272, filed on April 13, 2021; U.S. Provisional Patent Application No. 63 / 261,554, filed on September 23, 2021; U.S. Provisional Patent Application No. 63 / 264,831, filed on December 2, 2021; and U.S. Provisional Patent Application No. 63 / 362,799, filed on April 11, 2022, the contents of which are hereby incorporated by reference in their entirety for all purposes.
[0002] The present invention generally relates to a computer system having a display generation component and one or more input devices that present a graphical user interface, including but not limited to an electronic device that provides an immersive experience within a three - dimensional environment.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements such as buttons and other graphics.
[0004] However, the ways and interfaces for interacting with environments (such as applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are complex, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which the manipulation of virtual objects is complex and error-prone impose a large cognitive burden on the user and degrade the experience in the virtual / augmented reality environment. In addition, those ways are time-consuming more than necessary, thereby wasting energy. This latter consideration is particularly important in battery-operated devices. SUMMARY OF THE INVENTION
[0005] Accordingly, there is a need for a computer system having improved methods and interfaces for providing a computer-generated experience that makes interacting with the computer system more efficient and intuitive for the user. Such methods and interfaces can optionally complement or replace conventional ways of providing the user with a computer-generated reality experience. Such methods and interfaces reduce the number, degree, and / or type of inputs from the user by assisting the user in understanding the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.
[0006] The above-mentioned deficiencies and other problems related to the user interface for a computer system having a display generation component and one or more input devices are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generation component, and the output devices include one or more haptic output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or sets of instructions stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the user's body when captured by a camera and other motion sensors, and voice input when captured by one or more audio input devices.In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game play, making a phone call, video conferencing, sending an email, instant messaging, training support, digital photography, digital video shooting, web browsing, playing digital music, taking notes, and / or playing digital video. The executable instructions for performing those functions are optionally included in a non-transitory computer-readable storage medium or other computer program products configured to be executed by one or more processors.
[0007] There is a need for an improved method and interface for interacting with objects within a three-dimensional environment in an electronic device. Such a method and interface can complement or replace conventional methods for interacting with objects within a three-dimensional environment. Such a method and interface reduce the number, degree, and / or type of inputs from a user and generate a more efficient human-machine interface.
[0008] In some embodiments, the electronic device changes the immersion level of a virtual environment and / or spatial effects within the three-dimensional environment based on the geometric shape of the physical environment surrounding the device. In some embodiments, the electronic device modifies the virtual environment and / or spatial effects in response to detecting movement of the device. In some embodiments, the electronic device moves the user interface of an application inside and outside of the virtual environment. In some embodiments, the electronic device selectively changes the display of an imitation environment and / or atmosphere effects within the three-dimensional environment based on the movement of an object associated with the user's viewpoint. In some embodiments, the electronic device provides feedback to the user in response to the user moving a virtual object into and / or within the imitation environment according to some embodiments.
[0009] Note that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will be apparent to those skilled in the art, particularly in view of the drawings, the specification, and the claims. Further, note that the language used herein has been selected solely for readability and for the purpose of explanation, and not for the purpose of defining or limiting the subject matter of the invention.
[0010] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout the following figures.
Brief Description of the Drawings
[0011]
Figure 1
[0012]
Figure 2
[0013]
Figure 3
[0014]
Figure 4
[0015]
Figure 5
[0016]
Figure 6A
[0017]
Figure 6B
[0018]
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 7E
Figure 7F
Figure 7G
Figure 7H
[0019]
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 8E
Figure 8F
Figure 8G
Figure 8H
[0020]
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 9E
Figure 9F
Figure 9G
Figure 9H
[0021]
Figure 10A
Figure 10B
Figure 10C
Figure 10D
Figure 10E
Figure 10F
Figure 10G
Figure 10H
Figure 10I
Figure 10J
Figure 10K
Figure 10L
Figure 10M
Figure 10N
Figure 10O
Figure 10P
[0022]
Figure 11A
Figure 11B
Figure 11C
Figure 11D
[0023]
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 12E
Figure 12F
Figure 12G
[0024]
Figure 13A
Figure 13B
Figure 13C
Figure 13D
Figure 13E
Figure 13F
Figure 13G
Figure 13H
[0025]
Figure 14A
Figure 14B
Figure 14C
Figure 14D
Figure 14E
Figure 14F
Figure 14G
Figure 14H
Figure 14I
Figure 14J
Figure 14K
Figure 14L
[0026]
Figure 15A
Figure 15B
Figure 15C
Figure 15D
Figure 15E
Figure 15F
Figure 15G
[0027]
Figure 16A
Figure 16B
Figure 16C
Figure 16D
Figure 16E
Figure 16F
Figure 16G
Figure 16H
Best Mode for Carrying Out the Invention
[0028] The present disclosure relates to a user interface for providing a user with a computer-generated reality (CGR) experience according to some embodiments.
[0029] The systems, methods, and GUIs described herein provide an improved way for an electronic device to interact with and manipulate objects in a three-dimensional environment.
[0030] In some embodiments, a computer system displays a virtual environment in a three-dimensional environment. In some embodiments, the virtual environment is displayed via a far-field process or a near-field process based on the geometry (e.g., size and / or shape) of the three-dimensional environment (e.g., optionally simulating the real-world environment around the device). In some embodiments, the far-field process includes introducing the virtual environment from the location farthest from the user's viewpoint and gradually expanding the virtual environment towards the user's viewpoint (e.g., using information about the location, position, distance, etc. of objects within the environment). In some embodiments, the near-field process includes introducing the virtual environment from the location farthest from the user's viewpoint without considering the distance and / or position of objects within the environment and expanding it outward from the initial location (e.g., expanding the size of the virtual environment with respect to the display generation component).
[0031] In some embodiments, the computer system displays a virtual environment and / or an ambient effect within a three-dimensional environment. In some embodiments, displaying an ambient effect includes displaying one or more lighting effects and / or particle effects in the three-dimensional environment. In some embodiments, in response to detecting movement of the device, a portion of the virtual environment is de-emphasized, but optionally, the ambient effect is not reduced. In some embodiments, in response to detecting rotation of the user's body (e.g., simultaneously with rotation of the device), the virtual environment is optionally moved to a new location within the three-dimensional environment that is aligned with the user's body.
[0032] In some embodiments, the computer system displays a virtual environment simultaneously with the user interface of an application. In some embodiments, the user interface of the application is moved within the virtual environment and can be treated as a virtual object existing within the virtual environment. In some embodiments, the user interface is automatically resized when moved within the virtual environment based on the distance of the user interface when moved within the virtual environment. In some embodiments, while both the virtual environment and the user interface are being displayed, the user can request that the user interface be displayed as an immersive environment. In some embodiments, in response to a request to display the user interface as an immersive environment, the previously displayed virtual environment is replaced with the immersive environment of the user interface.
[0033] In some embodiments, the computer system displays a simulated environment and / or an ambient effect within a three-dimensional environment. In some embodiments, in response to detecting movement of a user, a part of the user (e.g., the head, eyes, face, and / or body), the computer system, and / or another component of the computer system, the simulated environment is moved back within the three-dimensional environment to expose a portion of the physical environment. In some embodiments, the ambient effect is not reduced. In some embodiments, further movement of a user, a part of the user (e.g., the head, eyes, face, and / or body), the computer system, and / or another component of the computer system causes the simulated environment to no longer be displayed.
[0034] In some embodiments, an electronic device displays a virtual object within a three-dimensional environment, the three-dimensional environment includes a simulated environment having a first visual appearance, and the virtual object is located outside the simulated environment within the three-dimensional environment. In some embodiments, in response to receiving an input to move the virtual object into or onto the simulated environment, the electronic device displays a simulated environment having a second visual appearance different from the first visual appearance. In some embodiments, changing the visual appearance of the simulated environment includes changing the location of the simulated environment, changing the immersion level associated with the simulated environment, changing the opacity of the simulated environment, and / or changing the color of the simulated environment.
[0035] The processes described below enhance the operability of a device and streamline the user interface with the device by various techniques, including other technologies, such as (for example, by assisting the user in making appropriate inputs when operating / interacting with the device and reducing user errors), providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional controls displayed, performing an operation without requiring further user input when a set of conditions is met, improving privacy and / or security, and enabling the user to use the device more quickly and efficiently, which reduces power usage and improves the battery life of the device.
[0036] Figures 1-6 provide an illustration of an exemplary computer system for providing a CGR experience (as described below with respect to methods 800, 1000, 1200, 1400, and 1600) to a user. In some embodiments, as shown in FIG. 1, the CGR experience is provided to the user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., speakers 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted device or a handheld device).
[0037] When describing the CGR experience, various related but distinct terms are used to refer individually to several environments with which the user can interact and / or that the user can perceive and / or (e.g., using inputs detected by the computer system 101 to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101 that generates the CGR experience) generate. The following is a subset of these terms.
[0038] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the aid of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through senses such as vision, touch, hearing, taste, and smell.
[0039] Computer-Generated Reality (or Extended Reality (XR)): In contrast, a computer-generated reality (CGR) environment refers to an environment that is wholly or partially simulated and with which people can perceive and / or interact through an electronic system. In CGR, a subset or representation of a person's body movements is tracked, and in response, one or more characteristics of one or more virtual objects simulated within the CGR environment are adjusted to behave according to at least one law of physics. For example, a CGR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to the person in a manner similar to how such views and sounds would change in the physical environment. Depending on the situation (e.g., for accessibility reasons), the adjustment of the characteristic(s) of the virtual object(s) in the CGR environment may be made in response to a representation of body movement (e.g., a voice command). A person may use any one of these senses, including vision, hearing, touch, taste, and smell, to perceive and / or interact with a CGR object. For example, a person can perceive and / or interact with an audio object that creates an audio environment with a 3D or spatial spread that provides the perception of a point sound source in 3D space. In another example, an audio object can enable audio transparency that selectively incorporates ambient sound from the physical environment, with or without including computer-generated audio. In some CGR environments, a person may only perceive and / or interact with audio objects.
[0040] Examples of CGR include virtual reality and mixed reality.
[0041] Virtual Reality: A virtual reality (VR) environment refers to an imitation environment designed to be based entirely on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of the person's presence within the computer-generated environment and / or through a simulation of a subset of the person's physical movements within the computer-generated environment.
[0042] Mixed Reality: In contrast to a VR environment designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment refers to an imitation environment designed to incorporate sensory inputs or representations thereof from the physical environment in addition to including computer-generated sensory inputs (e.g., virtual objects). On the virtual continuum, an MR environment is anywhere between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, the computer-generated sensory inputs can respond to changes in sensory inputs from the physical environment. Also, some electronic systems for presenting an MR environment may track the location and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical items or representations thereof from the physical environment). For example, the system can account for movement so that a virtual tree appears stationary relative to the physical ground.
[0043] Examples of mixed reality include augmented reality and augmented virtuality.
[0044] Augmented Reality: An augmented reality (AR) environment refers to an emulated environment in which one or more virtual objects are superimposed on a physical environment or its representation. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system synthesizes the image or video with virtual objects and presents the composite on the opaque display. A person uses this system to indirectly view the physical environment via the image or video of the physical environment and to perceive virtual objects superimposed on the physical environment. As used herein, a video of the physical environment shown on an opaque display is referred to as a "pass-through video," meaning that the system uses one or more image sensors (singular or plural) to capture an image of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into the physical environment or onto a physical surface, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. An augmented reality environment also refers to an emulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, when providing a pass-through video, the system may transform one or more sensor images to map to a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of the physical environment may be transformed by graphically modifying (e.g., magnifying) a portion thereof, whereby the modified portion can be made into a modified version that represents the original captured image but is non-photorealistic. As a further example, a representation of the physical environment may be transformed by graphically removing or obscuring a portion thereof.
[0045] Extended Virtual: An extended virtual (AV) environment refers to an emulated environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but people with faces are realistically reproduced from images of physical people. As another example, a virtual object may adopt the shape or color of a physical item imaged by one or more imaging sensors. As a further example, a virtual object can adopt a shadow that coincides with the position of the sun in the physical environment.
[0046] Viewpoint-locked virtual object: A virtual object is viewpoint-locked when the computer system displays the virtual object at the same location and / or position within the user's viewpoint even when the user's viewpoint shifts (e.g., changes). In an embodiment where the computer system is a head-mounted device, the user's viewpoint is locked in the forward direction of the user's head (e.g., the user's viewpoint is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's viewpoint remains fixed even when the user's line of sight moves without moving the user's head. In an embodiment where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generation component. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which the viewpoint-locked virtual object is displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an embodiment where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".
[0047] Environment Lock Virtual Object: A virtual object is environment locked (or "world locked") when a computer system displays the virtual object at a location and / or position within the user's perspective that is based on (e.g., selected with reference to and / or fixed to) a location and / or object within a three-dimensional environment (e.g., a physical environment or a virtual environment). When the user's perspective shifts, the location and / or object within the environment relative to the user's perspective changes, and as a result, the environment lock virtual object is displayed at a different location and / or position within the user's perspective. For example, an environment lock virtual object locked to a tree directly in front of the user is displayed at the center of the user's perspective. If the user's perspective shifts to the right (e.g., the user's head is turned to the right) and the tree moves to the left within the user's perspective (e.g., the position of the tree within the user's perspective shifts), the environment lock virtual object locked to the tree is displayed to the left within the user's perspective. In other words, the location and / or position at which the environment lock virtual object is displayed within the user's perspective depends on the location and / or position and / or orientation of the location and / or object within the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object within a physical environment) to determine the position at which to display the environment lock virtual object within the user's perspective. The environment lock virtual object can be locked to a stationary part of the environment (e.g., a floor, a wall, a table, or other stationary object), or to a movable part of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body such as the user's hand, wrist, arm, foot, etc. that moves independently of the user's perspective), such that the virtual object moves as the perspective or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0048] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits a delayed following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting the delayed following behavior, the computer system intentionally delays the movement of the virtual object when detecting movement of a reference point (e.g., a part of the environment, a viewpoint, or a point fixed relative to the viewpoint such as a point between 5 and 300 cm from the viewpoint) that the virtual object is following. For example, when the reference point (e.g., a part of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device so as to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops or decelerates, at which time the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits the delayed following behavior, the device ignores small movements of the reference point (e.g., movements of the reference point that are less than a threshold movement amount such as a movement of 0 to 5 degrees or a movement of 0 to 50 cm). For example, when the reference point (e.g., the part of the environment or the viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a position fixed or substantially fixed relative to a viewpoint or a part of the environment that is different from the reference point to which the virtual object is locked), and when the reference point (e.g., the part of the environment or the viewpoint to which the virtual object is locked) moves by a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed to maintain a position fixed or substantially fixed relative to a viewpoint or a part of the environment that is different from the reference point to which the virtual object is locked), and then decreases as the virtual object is moved by the computer system so as to maintain a position fixed or substantially fixed relative to the reference point, as the movement amount of the reference point increases beyond a threshold (e.g., a "delayed following" threshold).In some embodiments, the virtual object maintaining a position substantially fixed relative to the reference point includes a virtual object that is displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, or 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or front / back relative to the position of the reference point).
[0049] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display functionality, windows with integrated display functionality, displays formed as lenses designed to be placed on a person's eye (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers (singular or plural) and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing an image or video of the physical environment and / or one or more microphones for capturing the audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards the person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scan light source, or any combination of these technologies. The medium may be an optical waveguide, hologram medium, optical coupler, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The controller 110 will be described in more detail below with reference to FIG. 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the scene 105. In some embodiments, the controller 110 is communicatively coupled to a display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE802.11x, IEEE802.16x, IEEE802.3x, etc.). In another example, the controller 110 is included within the housing (e.g., the physical housing) of one or more of the display generation component 120 (e.g., an HMD, or a portable electronic device including a display and one or more processors, etc.), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0050] In some embodiments, the display generation component 120 is configured to provide the user with a CGR experience (e.g., at least the visual component of the CGR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The display generation component 120 will be described in more detail below with reference to FIG. 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0051] According to some embodiments, the display generation component 120 provides a CGR experience to the user while the user is virtually and / or physically present within the scene 105.
[0052] In some embodiments, the display generation component is worn on a part of the user's body (e.g., the user's own head or hand). Thus, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds a device with a display oriented towards the user's field of view and a camera oriented towards scene 105. In some embodiments, the handheld device is optionally disposed within a housing worn on the user's head. In some embodiments, the handheld device is optionally disposed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content while the user is not wearing or holding the display generation component 120. Many of the user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a device on a handheld or tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with CGR content triggered based on an interaction occurring within the space in front of a handheld or tripod-mounted device may be implemented in the same manner as an HMD where the interaction occurs within the space in front of the HMD and the response of the CGR content is displayed via the HMD. Similarly, a user interface showing an interaction with CRG content triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)) may be implemented in the same manner as an HMD where movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)).
[0053] Although the relevant features of the operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features for the sake of simplicity are not shown so as not to obscure more appropriate aspects of the exemplary embodiments disclosed herein.
[0054] FIG. 2 is a block diagram of an example of the controller 110 according to some embodiments. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more appropriate aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0055] In some embodiments, one or more communication buses 204 include circuitry that interconnects system components and controls communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.
[0056] The memory 220 includes a high-speed random access memory such as a dynamic random-access memory (DRAM), a static random-access memory (SRAM), a double-data-rate random-access memory (DDRRAM), or other random access solid-state memory devices. In some embodiments, the memory 220 includes non-volatile memory such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. The memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 220, or the non-transitory computer-readable storage medium of the memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and a CGR experience module 240.
[0057] The operating system 230 includes instructions for processing various basic system services and performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). To that end, in various embodiments, the CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.
[0058] In some embodiments, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of FIG. 1, and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To that end, in various embodiments, the data acquisition unit 242 includes instructions and / or logic therefor, and heuristics and metadata therefor.
[0059] In some embodiments, the tracking unit 244 is configured to map the scene 105 and track at least the position / location of the display generation component 120 with respect to the scene 105 of FIG. 1, and optionally with respect to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To that end, in various embodiments, the tracking unit 244 includes instructions and / or logic therefor, and heuristics and metadata therefor. In some embodiments, the tracking unit 244 includes the hand tracking unit 243 and / or the eye tracking unit 245. In some embodiments, the hand tracking unit 243 is configured to track the position / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand, with respect to the scene 105 of FIG. 1, with respect to the display generation component 120, and / or with respect to a coordinate system defined for the user's hand. The hand tracking unit 243 is described in more detail below with respect to FIG. 4. In some embodiments, the eye tracking unit 245 is configured to track the position and movement of the user's line of sight (or more generally the user's eyes, face, or head) with respect to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hand)), or with respect to the CGR content displayed via the display generation component 120. The eye tracking unit 245 is described in more detail below with respect to FIG. 5.
[0060] In some embodiments, the adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by the display generation component 120 and, optionally, by one or more of the output device 155 and / or the peripheral device 195. For that purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0061] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and, optionally, to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0062] Although the data acquisition unit 242, the tracking unit 244 (including, e.g., the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (including, e.g., the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 may be disposed within separate computing devices.
[0063] Furthermore, FIG. 2 is more intended to illustrate the functions of various features that may exist in a particular embodiment, as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules separately shown in FIG. 2 can be implemented in a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of particular functions and how functions are allocated among them, will vary depending on the implementation form, and in some embodiments, it depends in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation form.
[0064] FIG. 3 is a block diagram of an example of a display generation component 120 according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more suitable aspects of the embodiments disclosed herein. For that purpose, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional inward and / or outward image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0065] In some embodiments, one or more communication buses 304 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), accelerometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.), and the like.
[0066] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission device display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays, such as diffractive, reflective, polarizing, holographic, etc. For example, HMD 120 includes a single CGR display. In another example, HMD 120 includes a CGR display for each eye of the user. In some embodiments, one or more CGR displays 312 can present MR or VR content. In some embodiments, one or more CGR displays 312 can present MR or VR content.
[0067] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face, including the user's eyes (and may be referred to as an eye tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as a hand tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as a scene camera). The one or more optional image sensors 314 can include one or more RGB cameras, one or more infrared (IR) cameras, one or more event-based cameras, and / or the like (e.g., including a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor).
[0068] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or the non-transitory computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and a CGR presentation module 340.
[0069] The operating system 330 includes instructions for processing various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to a user via one or more CGR displays 312. For that purpose, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.
[0070] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of FIG. 1. For that purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0071] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. For that purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0072] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a composite reality scene or a map of a physical environment in which computer-generated objects can be placed) based on media content data. For that purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0073] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0074] Although the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 are shown as being present on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the CGR presentation unit 344, the CGR map generation unit 346, and the data transmission unit 348 may be arranged within separate computing devices.
[0075] Furthermore, FIG. 3 is more intended to illustrate the functions of various features that may exist in a particular implementation, as opposed to the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined, and some matters can be separated. For example, several functional modules separately shown in FIG. 3 can be implemented within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of particular functions and how functions are allocated among them, vary depending on the implementation and are partially dependent on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0076] FIG. 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (FIG. 1) is for the scene 105 of FIG. 1 (e.g., for a part of the physical environment surrounding the user, for the display generation component 120, or for a part of the user (e.g., the user's face, eyes, or head), and / or for a coordinate system defined with respect to the user's hand), the location / position of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand is tracked by the hand tracking unit 243 (FIG. 2). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0077] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a resolution sufficient to distinguish fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body or all of the body and can have either a zoom function or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.
[0078] In some embodiments, the image sensor 404 outputs a sequence of frames including 3D map data (and optionally color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API) and drives the display generation component 120 accordingly. For example, the user can interact with software running on the controller 110 by moving their hand 408 and changing the pose of their hand.
[0079] In some embodiments, the image sensor 404 projects a spot pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spots of the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a predetermined reference plane at a particular distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines a series of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the hand tracking device 440 can use other 3D mapping methods such as stereoscopy or time-of-flight measurement based on single or multiple cameras or other types of sensors.
[0080] In some embodiments, the hand tracking device 140 captures and processes a temporal sequence of depth maps of the user's hand while the user is moving their hand (e.g., the entire hand or one or more fingers). Software operating on a processor within the image sensor 404 and / or the controller 110 processes the 3D map data to extract hand patch descriptors within these depth maps. The software compares these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the hand pose in each frame. The pose typically includes the 3D locations of the user's hand joints and fingertips.
[0081] The software can also analyze the trajectories of the hand and / or fingers over multiple frames in a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, such that patch-based pose estimation is only performed once per two (or more) frames, while tracking is used to detect changes in pose occurring over the remaining frames. Pose, motion, and gesture information are provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 or perform other functions in response to pose and / or gesture information.
[0082] In some embodiments, the gestures include air gestures. An air gesture is a gesture that is detected without (or independently of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140), and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the other of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., tap gestures including movement of the hand in a predetermined pose by a predetermined amount and / or speed, or shake gestures including a predetermined speed or amount of rotation of a part of the user's body).
[0083] In some embodiments, the input gestures used in the various examples and embodiments described herein are air gestures implemented by movement of the user's finger(s) or other finger(s) or part(s) of the user's hand relative to other finger(s) or part(s) of the user's hand for interacting with a CGR or XR environment (e.g., a virtual or mixed reality environment) according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching (or independently of) an input element that is part of the device, and is based on detected movement of a part of the user's body, such as movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand from the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to the user's one hand, and / or movement of the user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined pose with a predetermined amount and / or speed, or a shake gesture including rotation of a part of the user's body at a predetermined speed or amount).
[0084] In some embodiments where the input gesture is an air gesture (i.e., there is no physical contact with an input device that provides information to the computer system regarding which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., line of sight) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is the detected attention (e.g., line of sight) of the user to a user interface element in combination with (e.g., simultaneously with) movement of the user's finger(s) and / or hand for performing pinch and / or tap inputs, as described in more detail below, for example.
[0085] In some embodiments, an input gesture directed to a user interface object is executed directly or indirectly with reference to the user interface object. For example, user input is executed directly on the user interface object in response to the user performing an input gesture with their hand at a position corresponding to the position of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current perspective). In some embodiments, the input gesture is executed indirectly on the user interface object in accordance with a user who performs the input gesture while the position of the user's hand is not at a position corresponding to the position of the user interface object in the three-dimensional environment, while detecting the user's attention (e.g., line of sight) to the user interface object. For example, in the case of a direct input gesture, the user can direct their input to the user interface object by starting a gesture at a position corresponding to or near the display position of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0 - 5 cm measured from the outer edge of the option or the central portion of the option). In the case of an indirect input gesture, the user can direct their input to the user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user starts an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position not corresponding to the display position of the user interface object).
[0086] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or augmented reality environment according to some embodiments. For example, the pinch inputs and tap inputs described later are performed as air gestures.
[0087] In some embodiments, the pinch input is part of an air gesture that includes one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, the pinch gesture, which is an air gesture, involves moving two or more fingers of a hand so as to touch each other, i.e., optionally including an interruption immediately after touching each other (e.g., within 0 to 1 second). The long pinch gesture, which is an air gesture, involves moving two or more fingers of a hand so as to touch each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption in the contact with each other. For example, the long pinch gesture includes the user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until an interruption in the contact between the two or more fingers is detected. In some embodiments, the double pinch gesture, which is an air gesture, includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected continuously and directly with each other (e.g., within a predetermined period). For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks the contact between two or more fingers), and then performs a second pinch input within a predetermined period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.
[0088] In some embodiments, a pinch-and-drag gesture, which is an air gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) that is performed in relation to (e.g., after) a drag input that changes the position of the user's hand from a first position (e.g., the starting position of the drag) to a second position (e.g., the ending position of the resistance). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opens two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches each other, and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture, which is an air gesture, includes an input (e.g., a pinch input and / or a tap input) that is performed using both hands of the user. For example, the input gesture includes two (e.g., or more) pinch inputs that are performed in relation to each other (e.g., simultaneously with a predetermined period or within a predetermined period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) performed using the user's first hand, and in relation to performing the pinch input using the first hand, a second pinch input is performed using the other hand (e.g., the second hand of the user's both hands). In some embodiments, movement between the user's both hands (e.g., to increase and / or decrease the distance or relative orientation between the user's both hands).
[0089] In some embodiments, a tap input that is performed as an air gesture (e.g., directed at a user interface element) includes a movement (s) of the user's finger (s) toward the user interface element, optionally a movement of the user's hand toward the user interface element with the user's finger (s) extended toward the user interface element, a downward movement of the user's finger (e.g., mimicking a mouse click action or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input that is performed as an air gesture is detected based on movement characteristics of a finger or hand that performs a tap gesture of the finger or hand away from the user's perspective and / or toward an object that is the target of the tap input where the end of the movement follows. In some embodiments, an end of the movement is detected based on a change in movement characteristics of a finger or hand that performs a tap gesture (e.g., movement away from the user's perspective and / or toward the object that is the target of the tap input, reversal of the direction of movement of the finger or hand, and / or reversal of the direction of acceleration of the movement of the finger or hand).
[0090] In some embodiments, the user's attention is determined to be directed at a portion of a three-dimensional environment based on detection of a line of sight directed at the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, in order for a device to determine that the user's attention is directed at a portion of a three-dimensional environment, while the user's perspective is within a distance threshold from the portion of the three-dimensional environment, at least a threshold duration (e.g., dwell time), a requirement that the line of sight be directed at the portion of the three-dimensional environment, and / or a requirement that the line of sight be directed at the portion of the three-dimensional environment, etc., based on detection of a line of sight directed at the portion of the three-dimensional environment with one or more additional conditions, the user's attention is determined to be directed at the portion of the three-dimensional environment, and if one of the additional conditions is not satisfied, the device determines that the attention is not directed at the portion of the three-dimensional environment where the line of sight is directed (e.g., until one or more additional conditions are satisfied).
[0091] In some embodiments, the detection of the readiness state configuration of the user or a part of the user is detected by the computer system. The detection of the hand readiness state configuration is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand. For example, the hand readiness state is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to perform a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's perspective (e.g., under the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular way (e.g., moved towards the area in front of the user above the user's waist and under the user's head, or away from the user's body or legs). In some embodiments, the readiness state is used to determine whether the interaction elements of the user interface respond to attention (e.g., gaze) input.
[0092] In some embodiments, the software may be downloaded in electronic form to the controller 110, e.g., over a network, or alternatively provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in FIG. 4 as a separate unit from the image sensor 440, by way of example, but some or all of the processing functions of the controller can be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the handtracking device 402, or otherwise. In some embodiments, at least some of these processing functions are performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, a handheld device, or a head-mounted device), or using any other suitable computerized device such as a game console or a media player. The sensing function of the image sensor 404 can similarly be integrated with a computer or other computerized device controlled by the sensor output.
[0093] FIG. 4 further includes a schematic diagram of a depth map 410 captured by an image sensor 404 according to some embodiments. The depth map includes a matrix of pixels each having a respective depth value. The pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The luminance of each pixel within the depth map 410 is inversely proportional to the depth value, i.e., the measured z - distance from the image sensor 404, and the tone becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment components (i.e., groups of adjacent pixels) of an image having the characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and movement from frame to frame of a sequence of depth maps.
[0094] FIG. 4 also schematically shows a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on the background 416 of the hand segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, the end of the hand connected to the wrist, etc.), and optionally major feature points on the wrist or arm connected to the hand are identified and placed on the hand skeleton 414. In some embodiments, the locations and movements of these major feature points over multiple image frames are used by the controller 110 to determine, according to some embodiments, a hand gesture or the current state of the hand being performed by the hand.
[0095] FIG. 5 shows an exemplary embodiment of the eye tracking device 130 (FIG. 1). In some embodiments, the eye tracking device 130 is controlled by an eye tracking unit 245 (FIG. 2) to track the position and movement of the user's line of sight with respect to scene 105 or CGR content displayed via display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed on a wearable frame, the head-mounted device includes both a component for generating CGR content for viewing by the user and a component for tracking the user's line of sight with respect to the CGR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, the eye tracking device 130 is optionally a device separate from the handheld device or the CGR chamber. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used with a display generation component worn on the head or a display generation component not worn on the head. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.
[0096] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that presents a frame including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and on which virtual objects can be displayed. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, whereby an individual can use the system to observe virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0097] As shown in FIG. 5, in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera), and an illumination source that emits light (e.g., IR or NIR light) towards the user's eyes (e.g., an IR or NIR light source such as an array or ring of LEDs). The eye tracking camera may be directed towards the user's eyes to directly receive reflected IR or NIR light from the light source, or alternatively, may be directed towards a "hot" mirror disposed between the user's eyes and a display panel that reflects IR or NIR light from the eyes while allowing visual light to pass through to the eye tracking camera. The gaze tracking device 130 optionally captures an image of the user's eyes (e.g., as a video stream captured at 60 - 120 frames per second (fps)), analyzes the image to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by an individual eye tracking camera and illumination source.
[0098] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility prior to delivery of the AR / VR device to the end user. The device-specific calibration process may be an automatic calibration process or a manual calibration process. The user-specific calibration process may include an estimation of the eye parameters of a particular user, such as pupil location, foveal location, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 are determined, the images captured by the eye tracking camera are processed using the glint assist method to determine the user's current visual axis and viewpoint with respect to the display.
[0099] As shown in FIG. 5, the eye tracking device 130 (e.g., 130A or 130B) includes a line-of-sight tracking system including one or more eyepieces 520, at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) disposed on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 is positioned between the user's eye(s) 592 and a display 510 (e.g., a display panel on the left or right side of a head-mounted display, or a display of a handheld device, a projector, etc.), and may be directed toward a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (as shown, for example, at the top of FIG. 5), or may be directed toward the user's eye(s) 592 to receive the reflected IR or NIR light from the eye(s) 592 (as shown, for example, at the bottom of FIG. 5).
[0100] In some embodiments, the controller 110 renders an AR or VR frame 562 (e.g., the left and right frames of the left and right display panels) and provides the frame 562 to the display 510. The controller 110 uses the eye tracking input 542 from the eye tracking camera 540, for example, when processing the frame 562 for display, for various purposes. The controller 110 optionally estimates the user's viewpoint on the display 510 based on the eye tracking input 542 obtained from the eye tracking camera 540 using a glint assist method or other suitable method. The viewpoint estimated from the eye tracking input 542 is optionally used to determine the direction the user is currently looking.
[0101] The following describes some possible use cases of the user's current line of sight direction, but this is not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined line of sight direction of the user. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current line of sight direction than in the peripheral region. As another example, the controller may position or move virtual content within the view at least partially based on the user's current line of sight direction. As another example, the controller may display specific virtual content within the view at least partially based on the user's current line of sight direction. As another exemplary use case in an AR application, the controller 110 can capture the physical environment of the CGR experience and direct the external camera to focus in the determined direction. The autofocus mechanism of the external camera can then focus on an object or surface within the environment that the user is currently viewing on the display 510. As another exemplary use case, the eyepiece 520 may be a focusable lens, and the line of sight tracking information is used by the controller to adjust the focus of the eyepiece 520 so that the virtual object the user is currently viewing has appropriate binocular convergence to match the convergence of the user's eyes 592. The controller 110 can utilize the line of sight tracking information to direct and adjust the focus of the eyepiece 520 so that the nearby object the user is viewing appears at the correct distance.
[0102] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye-tracking camera (e.g., eye-tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light source may be arranged in a ring or circularly around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be employed.
[0103] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, so as not to introduce noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as an example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0104] Embodiments of the eye-tracking system as shown in FIG. 5 can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide the user with an experience of computer-generated reality, virtual reality, augmented reality, and / or augmented virtuality.
[0105] FIG. 6A shows a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., the eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the glint-assisted gaze tracking system uses prior information from the previous frame when analyzing the current frame to track the pupil contour and glint within the current frame. If not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint within the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.
[0106] As shown in FIG. 6A, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0107] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, as shown at 620, the image is analyzed to detect the user's pupil and glint within the image. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If not successfully detected, the method returns to element 610 to process the next image of the user's eyes.
[0108] At 640, when proceeding from element 410, the current frame is analyzed to track the pupil and glint, based in part on previous information from the previous frame. At 640, when proceeding from element 630, the tracking state is initialized based on the detected pupil and glint within the current frame. The result of the processing at element 640 is checked to confirm that the result of the tracking or detection is reliable. For example, the result can be checked to determine whether a sufficient number of glints for performing pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the result is not reliable, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's viewpoint.
[0109] FIG. 6A is intended to function as an example of an eye tracking technique that can be used in a particular implementation. As will be recognized by those skilled in the art, other eye tracking techniques that currently exist or may be developed in the future can be used in computer system 101, instead of or in combination with the glint-assisted eye tracking technique described herein, to provide the user with a CGR experience according to various embodiments.
[0110] Figure 6B shows an exemplary environment of an electronic device 101 for providing a CGR experience according to some embodiments. In Figure 6B, the real-world environment 602 includes the electronic device 101, the user 608, and real-world objects (e.g., table 604). As shown in Figure 6B, the electronic device 101 is optionally attached to a tripod or otherwise fixed to the real-world environment 602 such that one or more hands of the user 608 are free (e.g., the user 608 is not optionally holding the device 101 with one or more hands). As described above, the device 101 optionally has one or more groups of sensors disposed on different sides of the device 101. For example, the device 101 optionally includes a sensor group 612-1 and a sensor group 612-2 located on the "back" side and the "front" side of the device 101, respectively (e.g., capable of capturing information from respective faces of the device 101). As used herein, the front side of the device 101 is the side facing the user 608, and the back side of the device 101 is the side facing away from the user 608.
[0111] In some embodiments, the sensor group 612-2 includes an eye-tracking unit (e.g., the eye-tracking unit 245 described above with reference to Figure 2) that includes one or more sensors for tracking the user's eyes and / or line of sight, and the eye-tracking unit can "look at" the user 608 and track the user's eye(s) in the manner described above. In some embodiments, the eye-tracking unit of the device 101 can capture the movement, orientation, and / or line of sight of the user 608's eyes and process the movement, orientation, and / or line of sight as an input.
[0112] In some embodiments, the sensor group 612-1 can include a hand tracking unit (e.g., the hand tracking unit 243 described above with reference to FIG. 2) that can track one or more hands of the user 608 held on the "back" side of the device 101, as shown in FIG. 6B. In some embodiments, the hand tracking unit can be optionally included in the sensor group 612-2 such that the user 608 can additionally or alternatively hold one or more hands on the "front" side of the device 101 while the device 101 tracks the position of one or more hands. As described above, the hand tracking unit of the device 101 can capture the movement, position, and / or gesture of one or more hands of the user 608 and process the movement, position, and / or gesture as an input.
[0113] In some embodiments, the sensor group 612-1 can optionally include one or more sensors (e.g., the image sensor 404 described above with reference to FIG. 4) configured to capture an image of the real-world environment 602 including the table 604. As described above, the device 101 can capture an image of a portion (e.g., part or all) of the real-world environment 602 and present the captured portion of the real-world environment 602 to the user via one or more display generation components of the device 101 (e.g., a display of the device 101 optionally located on the side of the device 101 facing the user, opposite to the side of the device 101 facing the captured portion of the real-world environment 602).
[0114] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with a CGR experience, e.g., a composite reality environment in which one or more virtual objects are superimposed on a representation of the real-world environment 602.
[0115] Accordingly, the description herein describes some embodiments of a three-dimensional environment (e.g., a CGR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists within a physical environment that is captured and displayed within the three-dimensional environment (e.g., actively via a camera and display of an electronic device or passively via a transparent or translucent display of the electronic device). As described above, the three-dimensional environment is optionally a mixed reality system based on a physical environment that is captured by one or more sensors of a device and displayed via a display generation component. As a mixed reality system, the device can optionally selectively display portions and / or objects of the physical environment such that each appears to exist within the three-dimensional environment displayed by the electronic device. Similarly, in the real world, the device can optionally display virtual objects within the three-dimensional environment such that the virtual objects appear to exist within the real world (e.g., the physical environment) by placing virtual objects at respective locations within the three-dimensional environment that have corresponding locations in the real world. For example, the device can optionally display a vase such that it appears as if a real vase is placed on a table within the physical environment. In some embodiments, each location within the three-dimensional environment has a corresponding location within the physical environment. Thus, when the device is described as displaying a virtual object at an individual location with respect to a physical object (e.g., the location of the user's hand or near it, or on or near a physical table, etc.), the device displays the virtual object at a particular location within the three-dimensional environment such that the virtual object appears to be on or near the physical object within the physical world (e.g., the virtual object is displayed at a location within the three-dimensional environment that corresponds to the location within the physical environment where the virtual object would be displayed if it were a real object at that particular location).
[0116] In some embodiments, real-world objects present in the physical environment that are displayed in the three-dimensional environment can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table in the physical environment and the vase is a virtual object.
[0117] Similarly, the user can optionally use one or more hands to interact with virtual objects in the three-dimensional environment as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the device optionally capture one or more of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or, in some embodiments, due to the transparency / semi-transparency of the user interface, or the projection of the user interface onto a transparent / semi-transparent surface, or the portion of the display generating component displaying the projection of the user interface onto the user's eye or field of view of the user's eye, the user's hands are visible through the display generating component with the ability to see the physical environment through the user interface. Thus, in some embodiments, the user's hands are displayed at separate locations in the three-dimensional environment and are processed as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, the user can move their hands to move a representation of the hands in the three-dimensional environment in coordination with the movement of the user's hands.
[0118] In some of the embodiments described below, for example, for the purpose of determining whether a physical object is interacting with a virtual object (e.g., whether a hand is touching, grasping, holding a virtual object, etc., or whether it is within a threshold distance from the virtual object), the device can optionally determine the "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the device determines the distance between the user's hand and the virtual object. In some embodiments, the device determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the target virtual object in the three-dimensional environment. For example, one or more hands of the user are placed at a specific position in the physical world, which the device can optionally capture and display at a specific corresponding position in the three-dimensional environment (e.g., the position in the three-dimensional environment where the hand is displayed if the hand is a virtual hand rather than a physical hand). The position of the hand in the three-dimensional environment is optionally compared with the position of the target virtual object in the three-dimensional environment to determine the distance between one or more hands of the user and the virtual object. In some embodiments, the device optionally determines the distance between a physical object and a virtual object by comparing positions in the physical world (as opposed to, for example, comparing positions in the three-dimensional environment). For example, when determining the distance between one or more hands of the user and a virtual object, the device optionally determines the corresponding location of the virtual object in the physical world (e.g., the position where the virtual object is located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and one or more hands of the user.In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether the physical object is within a threshold distance of the virtual object, the device optionally maps the location of the physical object into the three-dimensional environment and / or maps the location of the virtual object into the physical world by performing any of the techniques described above.
[0119] In some embodiments, the same or similar techniques are used to determine where and what the user's line of sight is directed to and / or where and what the physical stylus held by the user is directed to. For example, if the user's line of sight is directed to a particular position within the physical environment, the device optionally determines the corresponding position within the three-dimensional environment, and if a virtual object is located at that corresponding virtual position, the device optionally determines that the user's line of sight is directed to that virtual object. Similarly, the device can optionally determine where the stylus is pointing in the physical world based on the orientation of the physical stylus. In some embodiments, based on this determination, the device determines the corresponding virtual position within the three-dimensional environment corresponding to the location within the physical world that the stylus is pointing to and optionally determines that the stylus is pointing to the corresponding virtual position within the three-dimensional environment.
[0120] Similarly, the embodiments described herein may refer to the location of a user (e.g., a user of a device), and / or the location of a device within a three-dimensional environment. In some embodiments, the user of the device is holding, wearing, or otherwise located on or near the electronic device. Thus, in some embodiments, the location of the device is used as a proxy for the location of the user. In some embodiments, the location of the device and / or user within the physical environment corresponds to an individual location within the three-dimensional environment. In some embodiments, the individual location is the location where the "camera" or "view" of the three-dimensional environment extends. For example, if a user stands at a location facing an individual portion of the physical environment displayed by a display generation component, the location of the device is the location within the physical environment (and its corresponding location within the three-dimensional environment) where the user would view an object within the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the object was displayed by the display generation component of the device. Similarly, if a virtual object displayed within the three-dimensional environment were a physical object within the physical environment (e.g., the virtual object were located at the same location within the physical environment as within the three-dimensional environment and had the same size and orientation within the physical environment as within the three-dimensional environment), the location of the device and / or user is the position where the user would view the virtual object within the physical environment in the same position, orientation, and / or size (e.g., absolutely, and / or relative to each other, and in the real-world object) as the virtual object was displayed by the display generation component of the device.
[0121] In this disclosure, various input methods are described with respect to interaction with a computer system. If one example is provided using one input device or input method and another example is provided using another input device or input method, it should be understood that each example may be compatible with and optionally utilize the input device or input method described in another example. Similarly, various output methods are described with respect to interaction with a computer system. If one example is provided using one output device or output method and another example is provided using another output device or output method, it should be understood that each example may be compatible with and optionally utilize the output device or output method described in another example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. If one example is provided using an interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with and optionally utilize the method described in another example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.
[0122] Furthermore, in the methods described herein, which are conditional on one or more conditions being met by one or more steps, it should be understood that the methods can be repeated in multiple iterations such that, over the course of the repetition, all of the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires performing a first step when a condition is met and a second step when the condition is not met, one of ordinary skill in the art will understand that the steps recited in the claims will be repeated in a particular order until the condition is met and then ceases to be met. Thus, a method described in terms of one or more steps that are dependent on one or more conditions being met can be rewritten as a method that is repeated until each condition described in the method is met. However, this is not required in claims directed to a system or computer-readable medium that includes instructions for performing conditional operations based on the fulfillment of corresponding one or more conditions, and thus can determine whether an occurrence is satisfied without explicitly repeating the steps of the method until all conditions for the steps of the method being conditional are met. One of ordinary skill in the art will also understand that, similar to a method having conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as necessary to ensure that all of the conditional steps are executed. User Interface and Related Processes
[0123] Attention is now directed to embodiments of a user interface ("UI") and related processes that can be implemented in a computer system such as a portable multifunctional device or a head-mounted device, including a display generation component, one or more input devices, and (optionally) one or more cameras.
[0124] Figures 7A - 7H illustrate examples of displaying a virtual environment according to some embodiments.
[0125] FIG. 7A shows an electronic device 101 that displays a three-dimensional environment 704 on a user interface via a display generation component (e.g., the display generation component 120 of FIG. 1). As described above with reference to FIGS. 1-6, the electronic device 101 optionally includes a display generation component (e.g., a touch screen) and a plurality of image sensors (e.g., the image sensors 314 of FIG. 3). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the user interface shown below can also be implemented on a head-mounted display that includes a display generation component for presenting the user interface to the user, sensors for detecting the physical environment and / or movement of the user's hand (e.g., external sensors facing outward from the user), and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face). The figures in this specification show the three-dimensional environment presented to the user by the device 101 (e.g., and displayed by the display generation component of the device 101), and an overhead view of the physical environment and / or the three-dimensional environment associated with the device 101 for showing the relative locations of objects in the real-world environment and the locations of virtual objects in the three-dimensional environment (e.g., the overhead view 718 of FIG. 7A, etc.).
[0126] As shown in FIG. 7A, device 101 captures one or more images of the real-world environment 702 (e.g., operating environment 100) around device 101, including one or more objects within the real-world environment 702 around device 101. In some embodiments, device 101 displays a representation of the real-world environment within a three-dimensional environment 704. For example, the three-dimensional environment 704 includes a representation of the back corner of a room, a representation of a corner table 708a, a representation of a desk 710a, a representation of a picture frame 706 on the back wall of the room, a representation of a coffee table 714a, and a representation of a side table 712a. Thus, the three-dimensional environment 704 optionally reconstructs portions of the real-world environment 702 such that the three-dimensional environment 704 appears to the user as if the user were physically located within the real-world environment 702 (e.g., optionally, facing the direction the user is currently facing, from a perspective of the user's current location within the real-world environment 702).
[0127] As shown in the overhead view 718 of the real-world environment 702 of FIG. 7A, corner table 708a, desk 710b, side table 712b, and coffee table 714b are real objects within the real-world environment 702 captured by one or more sensors of device 101, and their representations are included in the three-dimensional environment 704 (e.g., realistic representation, simplified representation, cartoon, caricature, etc.). In some embodiments, the real-world environment 702 also includes a picture frame (e.g., represented within the three-dimensional environment 704 by representation 706) and a sofa 719.
[0128] As shown in FIG. 7A, user 720 of device 101 is sitting on sofa 719 and holding device 101 so as to face the other end of the room (e.g., such that one or more sensors face the other end of the room) (e.g., or, for example, if device 101 is a head-mounted device, wearing device 101), and thus capturing corner table 708a, desk 710b, side table 712b, and coffee table 714b and displaying a representation of the objects within three-dimensional environment 704. In some embodiments, device 101 is capable of determining the geometric shape of at least a portion of the real-world environment 702 (e.g., the portion of the real-world environment 702 that device 101 faces, all portions of the real-world environment 702, etc.). In some embodiments, device 101 uses one or more sensors such as visible light sensors (e.g., cameras), depth sensors (e.g., time-of-flight sensors), or a combination of sensors (e.g., multiple sensors of the same type, different types of sensors, etc.) to determine the geometric shape of the real-world environment 702. In FIG. 7A, since user 720 is sitting on sofa 719, device 101 determines that there is a threshold amount of physical space in front of device 101. For example, the distance from user 720 (e.g., and thus from device 101) to the back wall (where desk 710b is located) is greater than a threshold distance (e.g., greater than 3 feet, 6 feet, 10 feet, 30 feet, etc.). Thus, since three-dimensional environment 704 is at least a partial reproduction of real-world environment 702 as described above, three-dimensional environment 704 optionally reflects the available space between user 720 and the back wall of the real-world environment 702 (e.g., the distance from the perspective of three-dimensional environment 704 (e.g., the “camera” position from where the user views three-dimensional environment 704) to the back wall of three-dimensional environment 704 is the same or similar to the distance from user 720 to the back wall of the real-world environment 702), and three-dimensional environment 704 has at least a threshold amount of space (e.g., since the real-world environment 702 has at least a threshold amount of space in front of user 720).In some embodiments, as will be described in further detail below with respect to FIGS. 7C-7H, since there is at least a threshold amount of depth within the three-dimensional environment 704 in the front of the user's perspective view, the three-dimensional environment 704 is eligible for a long-distance field virtual environment transition animation.
[0129] In FIG. 7A, device 101 is displaying an immersion level indicator 716. In some embodiments, the immersion level indicator 716 indicates the current immersion level (e.g., of the maximum number of immersion levels) at which the device 101 is displayed in the three-dimensional environment 704. In some embodiments, the immersion level is the amount by which the view of the physical environment (e.g., the view of objects within the real-world environment 702) is obscured by a virtual environment (e.g., an emulated environment optionally different from the real-world environment 702 surrounding the user), or the amount by which objects in the physical environment are modified to achieve a particular spatial effect (e.g., as will be described in more detail below with respect to method 1000). For example, the maximum immersion level (e.g., full immersion) optionally refers to a state in which none of the physical environment is visible within the three-dimensional environment 704 via the display generation component 120 and the entire three-dimensional environment 704 is encompassed by the virtual environment. In some embodiments, an intermediate immersion level (e.g., an immersion level lower than the maximum and higher than no immersion) refers to a state in which a portion of the real-world environment 702 is visible within the three-dimensional environment 704 via the display generation component 120 and the portion of the real-world environment 702 that would otherwise be visible (e.g., if not for immersion) is replaced by the virtual environment. In some embodiments, the immersion level indicator 716 optionally includes a plurality of elements associated with a plurality of immersion levels. In some embodiments, the plurality of immersion levels are classified into secondary immersion levels and primary immersion levels, and the immersion level indicator 716 optionally includes tick marks (e.g., indicated by squares) and major tick marks (e.g., squares displayed with triangles). In some embodiments, as the immersion level increases, the device 101 presents more elements of the virtual environment.In some embodiments, the minor tick refers to an immersion state where a minor element is introduced or the size of a previously introduced element is increased (e.g., an increase in volume, an increase in the size of a virtual environment, etc.), and the major tick refers to an immersion state where, optionally in addition to increasing the size of a previously introduced element, a major element is introduced (e.g., a new visual element, a new audio element, etc.).
[0130] In FIG. 7A, the immersion level indicator 716 indicates that the current immersion level is no immersion (indicated by the fact that neither the square nor the triangle is filled), and thus, no virtual environment is displayed within the three-dimensional environment 704. In some embodiments, the immersion level indicator 716 is displayed when the current immersion level exceeds no immersion level and / or after a threshold time (e.g., 5 seconds, 10 seconds, 30 seconds, 1 minute, etc.) after the immersion level has changed and / or after a threshold time (e.g., 5 seconds, 10 seconds, 30 seconds, 1 minute, etc.) after the immersion level has decreased to zero.
[0131] FIG. 7B shows an embodiment similar to FIG. 7A, except that, as shown in the overhead view 718 of the real-world environment 702, user 720 is sitting at desk 710b. In some embodiments, the device 101 is located on the desk 710b and captures a view of the real-world environment 702 from the desk 710b to the back wall. Thus, the three-dimensional environment 704 is a view of a limited space from the user 720 to the back wall, and includes a representation of the back wall closer to the user, a representation 706 of a picture frame on the back wall, and a portion of the representation 710a of the desk 710b. In FIG. 7B, since the device 101 determines that there is no threshold amount of physical space in front of the device 101 in the room because the user 720 is sitting near the back wall of the room. In some embodiments, this determination by the device 101 is based on one or more of the distance from the device 101 to the back wall of the room, the area between the device 101 and the back wall of the room, or the volume between the device 101 and the back wall of the room. In some embodiments, this determination by the device 101 is additionally or alternatively based on the distance / area / volume within a predefined region relative to the user's perspective in the three-dimensional environment 704, or the average of the distance / area / volume of the nearest point, or the distance / area / volume of the farthest point.
[0132] For example, the distance from the user 720 to the back wall (and thus the distance of the device 101) is less than the threshold distance described above. Thus, since the three-dimensional environment 704 is at least a partial reproduction of the real-world environment 702 as described above, the three-dimensional environment 704 optionally reflects the available space between the user 720 and the back wall of the real-world environment 702 (e.g., the distance from the perspective of the three-dimensional environment 704 to the back wall of the three-dimensional environment 704 is the same or similar to the distance from the user 720 to the back wall of the real-world environment 702), and the three-dimensional environment 704 does not have at least a threshold amount of space (e.g., because the real-world environment 702 does not have at least a threshold amount of space in front of the user 720). In some embodiments, since there is no at least a threshold amount of depth in the three-dimensional environment 704 in front of the user's perspective view, the three-dimensional environment 704 is not eligible for a long-distance field virtual environment transition animation. Instead, as will be described in more detail below with respect to FIGS. 7C-7H, a short-distance field virtual environment transition animation will be performed.
[0133] In some embodiments, implementing a long - distance field transition animation when the real - world environment 702 has at least a threshold amount of space, and implementing a short - distance field transition animation when the real - world environment 702 does not have at least a threshold amount of space, provides an optimized experience that reduces the risk of dizziness or vertigo. For example, when a user is sitting in a location where there is a lot of space in front of the user, converting the space to a virtual environment with a potentially infinite horizon has a low risk of visual dissonance and dizziness. However, when the user is sitting in a location where there is little space in front of the user, converting that space to a virtual environment with a potentially infinite horizon has a high risk of visual dissonance and dizziness. Thus, implementing a long - distance field animation when the risk of dizziness is low provides a comfortable experience for the user, while implementing a short - distance field animation when the risk of dizziness is high allows the user to view the virtual environment, but slowly introduces the environment to the user to reduce the likelihood of dizziness.
[0134] Figures 7C and 7D each show a snapshot of a long - distance field transition animation and a short - distance field transition animation for displaying a virtual environment within a three - dimensional environment 704. In some embodiments, the virtual environment can be a simulated three - dimensional environment that is optionally instead of (e.g., for full immersion) or optionally simultaneously with (e.g., for partial immersion) a representation of the physical environment, displayed within the three - dimensional environment 704. Some examples of virtual environments include a lake environment, a mountain environment, a sunset scene, a sunrise scene, a nighttime environment, a grassland environment, a concert scene, etc. In some embodiments, the virtual environment is based on a real physical location such as a museum, an aquarium, etc. In some embodiments, the virtual environment is a location designed by an artist. Thus, displaying a virtual environment within the three - dimensional environment 704 provides the user with a virtual experience as if the user were physically located within the virtual environment.
[0135] In some embodiments, in response to user input selecting affordances associated with an individual virtual environment, the virtual environment is displayed in the three-dimensional environment 704. For example, the three-dimensional environment 704 includes one or more affordances of one or more virtual environments from which the user can select an individual environment for display in the three-dimensional environment 704. In some embodiments, in response to selecting an affordance for an individual environment, as described in more detail below with respect to method 1000, the individual environment is displayed by device 101 at a predetermined immersion level, such as a partial immersion level or a full immersion level. In some embodiments, the virtual environment is displayed within the three-dimensional environment 704 in response to user input that increases the immersion level of the three-dimensional environment 704. For example, device 101 optionally includes a mechanical dial that can be rotated in an individual direction to increase the immersion level and rotated in another direction to decrease the immersion level. Embodiments herein describe a process for displaying a virtual environment from a state where the virtual environment is not displayed to an intermediate immersion level or a full immersion level, and / or a process for changing the immersion level of a virtual environment from a non-immersion level or a partial immersion level (e.g., a level lower than full immersion) to a higher partial immersion level or a full immersion level.
[0136] Figures 7C-7D show a first snapshot of a process for displaying a virtual environment 722 at a high partial immersion level (e.g., as in Figures 7G-7H, where a large amount of the three-dimensional environment 704 is replaced by the virtual environment 722), although Figures 7C-7D optionally show snapshots of a process for displaying the virtual environment 722 at a low partial immersion level (e.g., where a small amount of the three-dimensional environment 704 is replaced by the virtual environment 722) and / or a full immersion level (e.g., where all of the three-dimensional environment 704 is replaced by the virtual environment 722). In some embodiments, as described above, the process for displaying the virtual environment 722 is triggered (e.g., started) in response to a user input that causes the virtual environment 722 to be displayed (e.g., where the virtual environment 722 was not displayed when the user input was received), or in response to a user input that changes the immersion level of the virtual environment 722 (e.g., where the virtual environment 722 was displayed when the user input was received).
[0137] FIG. 7C shows a first snapshot of the far-field process for displaying the virtual environment 722 when the immersion level is 1 (e.g., the level just above no immersion), as indicated by the immersion level indicator 716 where the leftmost square is filled and the other squares (and triangles) are not filled. As described above, in FIG. 7C, the process for displaying the virtual environment 722 includes a far-field transition, which involves replacing portions of the physical environment and / or the three-dimensional environment 704 with portions of the virtual environment based on the distance of those portions of the physical environment and / or the three-dimensional environment 704 from the user's viewpoint. For example, the farthest position in the three-dimensional environment 704 (e.g., the farthest position of the real-world environment 702 visible within the three-dimensional environment 704) is replaced with the virtual environment and expands from that farthest position towards the user's viewpoint (as shown, for example, in FIGS. 7E and 7G). In FIG. 7C, the left and right corners of the three-dimensional environment 704 correspond to the locations farthest from the user 720's viewpoint, and thus, the left corner of the three-dimensional environment 704 is replaced with the virtual environment 722-1a and the right corner of the three-dimensional environment 704 is replaced with the virtual environment 722-2a (as also shown, for example, by the virtual environment 722-1b and the virtual environment 722-2b in the overhead view 718). In some embodiments, the virtual environment 722-1a and the virtual environment 722-2a are part of the same virtual environment but correspond to the left and right sides (e.g., different portions) of the virtual environment 722. For example, the virtual environment 722 is at least partially superimposed on the three-dimensional environment 704 (e.g., optionally, the virtual environment 722 can cover all of the three-dimensional environment 704), but when the immersion level is zero (e.g., no immersion), the entire virtual environment 722 is hidden from view, and as the immersion level increases, portions of the virtual environment 722 are revealed (e.g., replacing and / or obscuring the view of each portion of the physical environment). Thus, in FIG. 7C, the left end portion of the virtual environment 722 is revealed so that the user can see the left end portion where the left rear corner was previously displayed, and the right end portion of the virtual environment 722 is revealed so that the user can see the right end portion of the virtual environment 722 where the right rear corner was previously displayed.
[0138] In some embodiments, the virtual environment 722 optionally has dimensions that are larger than the dimensions of the real-world environment 702 (e.g., if the virtual environment 722 is a landscape that extends to the horizon, e.g., potentially infinite depth), and thus, by replacing a portion of the three-dimensional environment 704 with the virtual environment 722, the replaced portion is displayed as if it were a visual portal into the virtual environment 722. For example, the overhead view 718 shows that the left and right rear corners of the real-world environment 702 no longer exist (e.g., are no longer displayed) within the three-dimensional environment 704 (e.g., as indicated by the dotted lines), and instead, optionally, extend within a distance according to the size, dimensions, and / or depth of the virtual environment 722 (e.g., extend outwardly from the left and right rear corners at a vertical angle from the user 720's perspective), a view into the virtual environment 722. Thus, in FIG. 7C, the three-dimensional environment 704 appears to the user as an environment that is partially the real-world environment 702 and partially the virtual environment 722, such that the user can enter the virtual environment 722 through the left and right rear corners of the real-world environment 702 as if a portion of the user's real-world environment 702 has been transformed into the virtual environment 722.
[0139] FIG. 7D shows a first snapshot of the near-field process for displaying a virtual environment 722 when the immersion level is 1 (e.g., the level just above no immersion), as indicated by the immersion level indicator 716 with the leftmost square filled and the other squares (and triangles) unfilled. As described above, in FIG. 7D, the process for displaying the virtual environment 722 optionally selects a central (or other fixed) location that optionally corresponds to a surface in the three-dimensional environment 704 that is farthest from the user (e.g., in the direction of the orientation of the user's viewpoint into the three-dimensional environment 704), and includes a near-field transition that includes expanding radially outward from the central location (e.g., optionally increasing the radius of the portal into the virtual environment regardless of the location and / or distance of the object in the three-dimensional environment 704 relative to the user 720, as will be described in more detail below with respect to FIGS. 7F and 7H). FIG. 7D shows a virtual environment 722a having a circular shape, but it should be understood that other shapes (e.g., oval, cylindrical, square, rectangular, etc.) are also possible. In FIG. 7D, the rear wall of the real-world environment 702 is the farthest surface within the three-dimensional environment 704 in the direction of the orientation of the user's viewpoint into the three-dimensional environment 704, and thus, starting from the central location of the rear wall, the device 101 replaces the central circular area of the rear wall with the virtual environment 722a (e.g., as also illustrated by the virtual environment 722b in the overhead view 718). In some embodiments, the z-position (e.g., depth) of the virtual environment 722a (e.g., the boundary between the virtual environment 722a and the real-world environment portion of the three-dimensional environment 704) is not necessarily the rear wall and is optionally a location closer to the user 720 than the rear wall. For example, in FIG. 7D, the overhead view 718 shows that the virtual environment 722b starts at the central depth location of the table 710b (e.g., there is a boundary between the real-world environment 702 and the virtual environment 722b). In FIG. 7D, the height of the virtual environment 722 coincides with the center of the surface of the rear wall (e.g., the boundary of the virtual environment 722 is directly in front of the rear wall), so the virtual environment 722 may appear to start from the surface of the rear wall.In FIG. 7D, the overhead view 718 shows that the center of the rear wall of the real-world environment 702 no longer exists (e.g., is no longer displayed) within the three-dimensional environment 704 (as indicated, for example, by the dotted line), and instead, optionally, is a view into the virtual environment 722 that extends within a distance according to the size, dimensions, and / or depth of the virtual environment 722 (e.g., extends outward from the rear wall at a perpendicular angle from the user 720's perspective).
[0140] As described above with respect to FIG. 7C, the virtual environment 722 is optionally superimposed on the three-dimensional environment 704, but is hidden (e.g., thus replacing or obscuring each part of the three-dimensional environment 704) until a part of the virtual environment 722 is exposed to the user according to the immersion level. Thus, in FIG. 7D, the device 101 exposes a part of the virtual environment 722 superimposed on the center of the rear wall (e.g., outward from the depth location of the center of the table 710b, based on the dimensions of the virtual environment 722). Thus, in FIG. 7D, the three-dimensional environment 704 appears to the user as an environment that is partially the real-world environment 702 and partially the virtual environment 722, as if a part of the user's real-world environment 702 has been converted to the virtual environment 722. Thus, even if both FIGS. 7C and 7D show the same immersion level in the transition for displaying the same virtual environment (e.g., the virtual environment 722), a part of the virtual environment 722 that is visible in FIG. 7D during the first snapshot of the near-field transition is different from a part of the virtual environment 722 that is visible in FIG. 7C during the first snapshot of the far-field transition.
[0141] Figures 7E - 7F show a second snapshot of the process of displaying the virtual environment 722 at a high immersion level (e.g., a continuation of the process described in FIGS. 7C - 7D, respectively, and ending at an immersion level as described in FIGS. 7G - 7H, respectively). As indicated by the immersion level indicator 716, the second snapshot corresponds to an immersion level of 3 and thus further increases the amount of the three - dimensional environment 704 replaced by the virtual environment 722a, and thus increases the size of the virtual environment 722a (e.g., increases the size of the visual "portal" into the virtual environment 722a).
[0142] For example, in FIG. 7E (which is, for example, a snapshot of a long-distance field transition animation), the virtual environment 722a expands towards the user in response to an increase in the immersion level (e.g., based on the distance of the physical environment and / or those parts of the three-dimensional environment 704 from the user's perspective) (e.g., replacing more of the real-world environment part of the three-dimensional environment 704). In some embodiments, expanding the virtual environment 722a includes expanding and / or moving the boundary between the virtual environment 722a and the real-world environment part of the three-dimensional environment 704 towards the user (e.g., as also shown in the top-down view 718). In some embodiments, the boundary of the virtual environment 722a remains equidistant from the user 720 and thus appears to move towards the user in a circular and / or spherical shape (when considered in three dimensions). In some embodiments, other shapes are possible for the expansion of the boundary of the virtual environment 722a. For example, the virtual environment 722a expands as a plane towards the user (e.g., a planar boundary parallel to the user that moves towards the user as the virtual environment 722a expands). In some embodiments, when the virtual environment 722a expands, parts of the real-world environment part of the three-dimensional environment 704 are no longer displayed and are replaced by parts of the virtual environment 722a that were not previously displayed. For example, in FIG. 7E, the back wall, the corner table 708a, and a part of the desk 710a (e.g., the more distant part) are no longer visible because they are farther from the user 720 than the boundary of the virtual environment 722a. In some embodiments, objects within the real-world environment part of the three-dimensional environment that are farther from the boundary of the virtual environment 722a (e.g., farther from the user's perspective) are no longer displayed, and objects within the real-world environment part of the three-dimensional environment that are closer to the boundary of the virtual environment 722a (e.g., closer to the user's perspective) continue to be displayed. In some embodiments, if an object has a height that obscures a part of the virtual environment 722a that would otherwise be part of it (e.g., the desk 710a in FIG. 7E), the object is optionally removed from the three-dimensional environment 704.For example, the right rear corner of the side table 712a obscures a portion of the virtual environment 722a in FIG. 7E and is optionally removed (e.g., no longer displayed) from the three-dimensional environment 704. In some embodiments, objects that are closer to the boundary of the virtual environment 722a but otherwise obscure a portion of the virtual environment 722a continue to be displayed (e.g., the side table 712a in FIG. 7E). In some embodiments, if an object straddles the boundary of the virtual environment 722a (e.g., the desk 710a in FIG. 7E, where a portion of the object is closer to the boundary and a portion is farther from the boundary), the object is optionally removed from the display (e.g., removed from the three-dimensional environment 704 so that it is no longer visible). In some embodiments, if an object straddles the boundary, a portion of the object closer to the boundary continues to be displayed, but a portion of the object farther from the boundary is no longer displayed (e.g., the desk 710a in FIG. 7E). In some embodiments, if an object straddles the boundary, the entire object continues to be displayed such that a portion of the object exists in the real-world portion of the three-dimensional environment 704 and a portion of the object exists in the mimicked environment portion of the three-dimensional environment 704 (e.g., optionally visually de-emphasized such as grayed out, darkened, made partially transparent, displayed as an outline, as if that portion of the object is part of the mimicked environment). Thus, in some embodiments, since at least a portion of the desk 710a is farther from the boundary of the virtual environment 722a, the entire desk 710a is no longer displayed (e.g., removed from the display). Alternatively, in some embodiments, the entire desk 710a continues to be displayed even if a portion of the desk 710a is farther from the boundary of the virtual environment 722a (e.g., such that the desk 710a appears to be partially within the virtual environment 722a and partially within the real-world environment portion of the three-dimensional environment 704).
[0143] In some embodiments, expanding the virtual environment 722a includes making more of the virtual environment 722a visible. For example, if the boundary of the virtual environment 722a is 10 feet away from the user, the user can see a portion of the virtual environment 722a 10 feet ahead, but if the boundary of the virtual environment 722a is expanded to be 5 feet away from the user, the user can see a portion of the virtual environment 722a 5 feet ahead that includes a portion of the virtual environment 722a 10 feet ahead (e.g., the portion shown when the boundary was 10 feet away) and an additional 5 feet of space that was not previously visible. Thus, as the virtual environment 722a expands, more of the virtual environment becomes visible to the user. For example, in FIG. 7E, the area between two corners of a room (e.g., the area between 722-1a and 722-2a in FIG. 7C that previously included a representation of the physical environment) is now occupied by the virtual environment 722a. In some embodiments, expanding the virtual environment 722a does not include moving the virtual environment 722a closer to the user. For example, an object within the virtual environment 722a that was previously 20 feet away from the user does not optionally move closer to the user as the boundary of the virtual environment 722a moves towards the user; rather, it remains 20 feet away. Instead, for example, an object within the virtual environment 722a that was 10 feet away and was not previously exposed to the user (e.g., due to the boundary being more than 10 feet away) is now optionally exposed to the user (e.g., due to the boundary moving closer to less than 10 feet away). Thus, expanding the virtual environment 722a causes the visual portal into the virtual environment 722a to expand and / or appear to move closer to the user without the features of the virtual environment 722a (e.g., the objects therein) moving closer to the user.
[0144] In FIG. 7F (e.g., a snapshot of a near-field transition animation), the virtual environment 722a optionally extends radially outward from a central location (e.g., the position shown above in FIG. 7D), regardless of the position and / or distance of the objects in the three-dimensional environment 704 relative to the user 720 (e.g., replacing more of the real-world environment portion of the three-dimensional environment 704). In some embodiments, radially expanding the virtual environment 722a optionally includes expanding the size / portal of the virtual environment 722a without moving the boundary of the virtual environment 722a closer to the user. For example, in FIG. 7F, the virtual environment 722b has a 30-degree radial size (e.g., the angle formed by the left and right boundaries of the virtual environment 722b is 30 degrees, and optionally, the angle formed by the bottom and top boundaries of the virtual environment 722b is 30 degrees), which is an increase in radial size compared to FIG. 7D. In some embodiments, the boundary of the virtual environment 722a expands in a circular fashion (or spherical fashion when considered in three dimensions), curving around the user 720, as shown by the virtual environment 722b in the overhead view 718. In some embodiments, the boundary of the virtual environment 722a expands outward and optionally remains equidistant from the user 720 along the expanded boundary (e.g., as if the boundary of the virtual environment 722a expands along the surface of a circle / sphere centered on the user 720). Thus, as shown in FIG. 7F, the virtual environment 722a is a circular visual portal that not only increases in radius (e.g., expands in the left-right and up-down directions parallel to the plane of the rear wall) but also curves around the user (e.g., changes in depth away from the plane of the rear wall). As described above, by expanding the virtual environment 722a, more of the virtual environment 722a becomes visible, replacing portions of the real-world environment. In FIG. 7F, a larger area of the rear wall is no longer visible, replaced by a portion of the virtual environment 722a that was not previously shown (e.g., compared to FIG. 7D), and a portion of the desk 710a is no longer visible.In some embodiments, the curvature of the virtual environment 722a around the user (e.g., in the z-dimension) for the near-field transition animation (e.g., FIGS. 7D, 7F, and 7H) is the same as or similar to the curvature of the virtual environment 722a (e.g., in the z-dimension) for the far-field transition animation (e.g., FIGS. 7C, 7E, and 7G) (e.g., same radius, same shape, etc.).
[0145] FIGS. 7G-7H show a third snapshot of the process of displaying the virtual environment 722 at a high immersion level (e.g., optionally, the final state, the stable state, the state of the three-dimensional environment 704 at the end of the process of displaying the virtual environment 722 including the first and second snapshots described above with respect to FIGS. 7C-7D and FIGS. 7E-7F, respectively). As indicated by the immersion level indicator 716, the third snapshot corresponds to an immersion level of 6, and thus further increases the amount of the three-dimensional environment 704 replaced by the virtual environment 722a, and thus increases the size of the virtual environment 722a (e.g., increases the size of the visual "portal" to the virtual environment 722a).
[0146] As shown in FIG. 7G showing a long-distance field transition animation, the virtual environment 722a is further expanded such that the boundary of the virtual environment 722a moves closer to the user 720 (e.g., optionally, while remaining equidistant from the user 720 along the boundary of the virtual environment 722a). For example, a portion of the three-dimensional environment between the representation 714a of the coffee table 714b and the representation 710a of the desk, which was not previously occupied by the virtual environment 722a, is now occupied by the virtual environment 722a. As described above, expanding the boundary of the virtual environment 722a towards the user does not involve moving the objects within the virtual environment 722a, but rather involves exposing (e.g., unmasking) a portion of the virtual environment 722a that is located closer to the user's perspective but was previously hidden. In some embodiments, since the virtual environment 722a encompasses more of the three-dimensional environment 704, more representations of physical objects extend beyond the boundary of the virtual environment 722a and are no longer displayed. For example, in FIG. 7G, the representation 710a of the desk 710b and the representation 712a of the side table 712b are no longer displayed, but the representation 714a of the coffee table 714b remains displayed because it is not located farther away than the boundary of the virtual environment 722a. Thus, the physical environment is no longer visible beyond the boundary of the virtual environment 722a (e.g., where only the virtual environment 722a exists), but the physical environment closer to the user than the boundary of the virtual environment 722a remains visible.
[0147] In FIG. 7H showing the near-field transition animation, the virtual environment 722a is further radially outwardly extended compared to the virtual environment 722a of FIG. 7F. For example, in FIG. 7H, the virtual environment 722b has a radial size of 180 degrees (e.g., the left and right boundaries of the virtual environment 722a extend immediately to the left and right of the user 720, and optionally, the upper and lower boundaries of the virtual environment 722a extend immediately above and below the user 720), which is an increase in the radial size compared to FIG. 7F, and thus, more of the desk 710a that was not previously occupied by the virtual environment 722a is now occupied by the virtual environment 722a. As described above, the virtual environment 722a optionally expands radially as if along the surface of a sphere surrounding the user 720. Thus, as the virtual environment 722a expands, the virtual environment 722a appears as a circle that expands (e.g., in the x-y dimension), but as shown in the top view 718, it also encloses the user (e.g., in the z dimension). In some embodiments, the virtual environment 722a does not expand closer to the user 720, and thus, a portion of the desk 710b is still visible at the bottom of the three-dimensional environment 704. As described above, the shape of the boundary of the virtual environment 722a is optionally the same as that of FIG. 7G (e.g., far-field transition animation) in FIG. 7H (e.g., near-field transition animation).
[0148] In some embodiments, FIGS. 7G and 7H show animations for displaying the virtual environment 722a and / or the final state of the display of the virtual environment 722a at a relatively high immersion level (e.g., the maximum immersion level, or an immersion level greater than the medium immersion level), and the amount of the three-dimensional environment 704 consumed by the virtual environment 722a is the same regardless of whether the animation is a long-distance field transition animation or a short-distance field transition animation. The size and shape of the virtual environment 722a in FIG. 7G (e.g., long-distance field transition animation) are optionally the same as the size and shape of the virtual environment 722a in FIG. 7H (e.g., short-distance field transition animation). For example, the boundary of the virtual environment 722a has the same shape between FIGS. 7G and 7H, and optionally, is at the same distance from the user's perspective between FIGS. 7G and 7H. In some embodiments, the user's view of the virtual environment 722a in FIGS. 7G and 7H is the same (e.g., individual objects within the virtual environment 722a are at the same distance and the same location relative to the user in FIGS. 7G and 7H).
[0149] Accordingly, in some embodiments, at a certain immersion level(s) (e.g., the maximum immersion level, a high immersion level, etc.), the virtual environment appears the same (e.g., the same size, the same shape, the same amount of the three-dimensional environment occupied by the virtual environment, the same portion of the virtual environment, etc.) to the user regardless of the amount of space in front of the user within the user's physical environment. In some embodiments, at a particular immersion level (e.g., the medium immersion level, the low immersion level, etc.), the virtual environment appears different based on whether there is sufficient space in front of the user or not for utilizing a long-distance field transition animation. For example, in the embodiments illustrated in FIGS. 7G-7H, at an immersion level higher than 6, the size and / or shape of the entrance into the virtual environment 722a are the same or similar between long-distance field transition and short-distance field transition, but at an immersion level lower than 6 (e.g., as described with reference to FIGS. 7C-7F), the size and / or shape of the entrance into the virtual environment 722a are different between long-distance field transition and short-distance field transition.
[0150] Figures 8A - 8H are flowcharts showing a method 800 for displaying a virtual environment according to some embodiments. In some embodiments, method 800 is implemented in a computer system (such as the computer system 101 of FIG. 1, such as a tablet, smartphone, wearable computer, or head - mounted device) that includes a display generation component (such as the display generation component 120 of FIGS. 1, 3, and 4) (such as a head - up display, display, touch screen, projector, etc.) and one or more cameras (such as a camera facing down with the user's hand (such as a color sensor, infrared sensor, and other depth - sensing cameras), or a camera facing forward from the user's head). In some embodiments, method 800 is stored in a non - transitory computer - readable storage medium and is controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (such as the control unit 110 of FIG. 1A). Some operations of method 800 may optionally be combined and / or the order of some operations may optionally be changed.
[0151] In method 800, in some embodiments, an electronic device (e.g., computer system 101 of FIG. 1) that communicates with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer) receives, via one or more input devices, an input corresponding to a request to display a first simulated environment, such as an affordance while displaying a three-dimensional environment 704 such as any of FIGS. 7A-7B or a selection of a representation of a simulated environment, while displaying an individual environment (e.g., a representation of a physical environment, a simulated environment, computer-generated reality, extended reality, etc.) via the display generation component (802) (e.g., detecting, via a hand-tracking device, an eye-tracking device, or any other input device, one or more hand movements of the user, one or more eye movements of the user, and / or any other input corresponding to a request to display the first simulated environment).
[0152] In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touch screen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) that projects a user interface and makes the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touch screen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, and / or a motion sensor (e.g., a hand-tracking sensor, a hand motion sensor).
[0153] In some embodiments, the individual environment is an augmented reality environment or a mixed reality environment that optionally includes virtual objects and / or optionally includes representations of real-world objects in the physical world surrounding the electronic device. In some embodiments, the individual environment is at least based on the physical environment surrounding the electronic device. For example, the electronic device can capture visual information about the environment surrounding the user (e.g., objects in the environment, the size and shape of the environment, etc.) and display at least a portion of the physical environment surrounding the user to the user, optionally making it appear as if the user is still located within the physical environment. In some embodiments, the individual environment (e.g., the physical environment surrounding the electronic device, a portion of the physical environment surrounding the electronic device, etc.) is actively displayed to the user via a display generation component. In some embodiments, the individual environment (e.g., the physical environment surrounding the electronic device, a portion of the physical environment surrounding the electronic device, etc.) is passively presented to the user via a partially transparent or translucent display through which the user can see at least a portion of the physical environment.
[0154] In some embodiments, the request to display the first simulated environment includes the operation of a rotational element, such as a mechanical dial or a virtual dial, of the electronic device or in communication with the electronic device. In some embodiments, the request includes the selection of displayed affordances and / or the operation of displayed control elements to increase the immersion level of the device and / or the individual environment. In some embodiments, the request includes a predetermined gesture recognized as a request to increase the immersion level of the device and / or the individual environment. In some embodiments, the request includes voice input that requests an increase in the immersion level of the device and / or the individual environment and / or requests the display of the first simulated environment. In some embodiments, the first simulated environment includes a scene that at least partially covers at least a portion of the individual environment such that the user appears to be located within the scene (e.g., optionally, no longer located within the first simulated environment). In some embodiments, the first simulated environment is an atmosphere transformation that modifies one or more visual characteristics of the individual environment such that the individual environment appears to be located at a different time, place, and / or condition (e.g., morning lighting instead of afternoon lighting, sunny instead of cloudy). In some embodiments, the immersion level includes the degree to which the content displayed by the electronic device (e.g., virtual environment) obscures the background content around / behind the virtual environment (e.g., content outside the virtual environment), optionally including the number of items of background content displayed, and the visual characteristics (e.g., color, contrast, opacity) in which the background content is displayed, and / or the angular range of the content displayed via the display generation component (e.g., 60 degrees for low-immersion displayed content, 120 degrees for medium-immersion displayed content, 180 degrees for high-immersion displayed content), and / or the percentage of the field of view consumed by the display generation consumed by the virtual environment (e.g., 33% of the field of view consumed by the virtual environment for low immersion, 66% of the field of view consumed by the virtual environment for medium immersion, 100% of the field of view consumed by the virtual environment for high immersion). In some embodiments, the background content is included in the background in which the virtual environment is displayed.In some embodiments, the background content includes a user interface (e.g., a user interface generated by a device corresponding to an application), virtual objects not associated with or not included in the virtual environment (e.g., files generated by a device, representations of other users, etc.), and / or real objects (e.g., pass-through objects representing real objects in the physical environment around the user's viewpoint that are visible through a display generation component so that the electronic device does not obscure / prevent their visibility through the display generation component and / or is visible through a transparent or translucent display generation component). In some embodiments, at a first (e.g., low) immersion level, the background, virtual, and / or real objects are displayed in an obscured manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with the background content, and the background content is optionally displayed with full brightness, color, and / or translucency. In some embodiments, at a second (e.g., high) immersion level, the background, virtual, and / or real objects are displayed in an obscured manner (e.g., dimmed, blurred, removed from the display, etc.). For example, an individual virtual environment with a high immersion level is displayed without simultaneously displaying the background content (e.g., in full screen or full immersion mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with the background content that is darkened, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of the background objects vary among the background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) compared to one or more second background objects, and one or more third background objects cease to be displayed.
[0155] In some embodiments, in response to receiving an input corresponding to a request to display a first emulation environment, the electronic device transitions from the display of an individual environment, such as from the display of the three-dimensional environment 704 without the virtual environment of FIGS. 7A-7B to the display of the virtual environment 722 of FIGS. 7G-7H, to the display of the first emulation environment (804) (e.g., causes the first emulation environment to be displayed in a portion of the display area (e.g., a subset of the display area, all of the display area, etc.)). In some embodiments, a portion of the individual environment is no longer displayed and / or is no longer visible to the user, and a portion of the first emulation environment is displayed in its place and / or is visible. In some embodiments, a portion of the individual environment is overlaid by the first emulation environment. In some embodiments, a portion of the individual environment is replaced by the first emulation environment.
[0156] In some embodiments, transitioning involves, in accordance with a determination that the physical environment surrounding the user's perspective meets one or more first criteria, transitioning from displaying an individual environment to displaying a first mimicked environment using a first type of transition, such as the long - range field transitions shown in FIGS. 7C, 7E, and 7G (806) (e.g., if the physical environment has a size, area, depth (e.g., in front of the electronic device) that exceeds a threshold, the shape of the physical environment conforms to a specific pattern, and / or a portion of the physical environment relative to the device has a size, area, depth (e.g., the space in front of the device, the space to the left of the device, the space to the right of the device, etc.) that exceeds a threshold, the transition from displaying an individual environment to displaying a first mimicked environment is a first type of transition), and, in accordance with a determination that the physical environment surrounding the user's perspective meets one or more second criteria that are different from the one or more first criteria, transitioning from displaying each environment to displaying a first mimicked environment using a second type of transition that is different from the first type of transition, such as the short - range field transitions shown in FIGS. 7D, 7F, and 7H (808) (e.g., if the physical environment surrounding the user's perspective has a size, area, depth (e.g., in front of the electronic device) that is less than a threshold, the shape of the physical environment conforms to a different pattern, and / or a portion of the physical environment relative to the device has a size, area, depth (e.g., the space in front of the device, the space to the left of the device, the space to the right of the device, etc.) that is less than a threshold, the transition from displaying each environment to displaying a first mimicked environment is a second type of transition), and includes.
[0157] Accordingly, one or more first criteria include criteria that are met when the physical environment is larger than a threshold size (e.g., 20 square feet, 50 square feet, 100 square feet, 300 square feet, etc.), area, and / or depth (e.g., the depth of the physical environment in front of the user and / or device has a depth greater than 2 feet, 5 feet, 8 feet, 10 feet, etc.). In some embodiments, the depth of the physical environment is determined based on the distance to the closest physical object (e.g., optionally, the closest physical object other than the floor, such as a coffee table or desk, in front of the user and / or device, or to the left, right, or behind the user and / or device). In some embodiments, the depth of the physical environment is determined based on the distance from a vertical plane such as a wall (e.g., optionally the closest vertical plane, or optionally the closest vertical plane in front of the user or device). In some embodiments, the depth of the physical environment is determined based on the distance from the closest boundary of the physical environment (or the boundary of the physical environment directly in front of the user or device). As described above, an individual environment may be at least partially based on the physical environment. Accordingly, in some embodiments, the characteristics of the physical environment notify of a transition between the individual environment and the first mimicked environment, for example, reducing transition shock, dizziness, and / or cognitive dissonance. For example, if the physical environment around and / or in front of the user / electronic device has a maximum depth exceeding 4 feet (e.g., the user is sitting at least 4 feet away from the wall opposite the user), the transition to the first mimicked environment includes gradually transitioning portions of the individual environment based on the distance from the electronic device and / or the user. For example, the area of the individual environment farthest from the user transitions first, then the next closest area transitions second, and so on, optionally continuing until reaching the user and / or the electronic device, or a predetermined distance in front of the user and / or the electronic device (e.g., 6 inches in front of the user, 1 foot in front of the user, 3 feet in front of the user, etc.).In some embodiments, different portions of the individual environment transition at different times (e.g., closer portions transition after farther portions). In some embodiments, non-adjacent areas of the individual environment transition simultaneously (e.g., different portions of the individual environment are not adjacent but are equidistant from the user and / or the electronic device). In some embodiments, the first type of transition starts from the farthest location and expands towards the user (e.g., as a plane moving towards the user or circularly towards the user, e.g., such that the boundary of the simulated environment is optionally equidistant from the user and / or the device throughout the transition). Thus, in some embodiments, the first type of transition depends on the difference in depth of objects within the individual environment (e.g., the order of which transitions are first, second, third, last, etc. is based on depth information). In some embodiments, transitioning includes dissolving individual portions of the individual environment into individual portions of the first simulated environment. In some embodiments, other transition animations are possible. In some embodiments, the portion of the individual environment that transitions to a portion of the first simulated environment is added to the previously transitioned portion of the first simulated environment. For example, the first portion of the individual environment transitions to the first portion of the first simulated environment, and then the second portion of the individual environment transitions to the second portion of the first simulated environment, and the first and second portions of the first simulated environment together form a continuous simulated environment.
[0158] In some embodiments, the user's perspective is a location within a physical environment (e.g., and / or its corresponding location within the three-dimensional environment, depending on the context of its use) having a perspective view (e.g., a view) of the physical environment that is the same (or similar) to the perspective view of the three-dimensional environment. For example, the three-dimensional environment is constructed to appear to the user as if the user were physically located at an individual location within the real-world environment (e.g., the representation of real-world objects within the three-dimensional environment appears to be the same distance away from the user at the same relative location as if the user were actually looking at the real-world objects within the real-world environment). Thus, one or more second criteria include criteria that are met when the physical environment has a threshold size (e.g., 20 square feet, 50 square feet, 100 square feet, 300 square feet, etc.), area, and / or depth (e.g., the space in front of the user and / or the device has a depth of less than 2 feet, 5 feet, 8 feet, 10 feet, etc.). In some embodiments, the second type of transition includes starting the transition from an individual location within an individual environment and expanding outward from the individual location. For example, if the physical environment around and / or in front of the user has a maximum depth of less than 4 feet, starting from a portion of the individual environment that is farthest from the user (e.g., the farthest point within the physical environment), the individual environment transitions to a first mimicked environment, and the transition expands outward (e.g., radially) from the starting position (e.g., with the starting position as the center of the transition, more of the individual environment transitions to the first mimicked environment). In some embodiments, the second type of transition starts from a specific location within an individual environment and expands outward from that location (e.g., left and right, up and down), and optionally, expands around the user (e.g., a circle that expands as opposed to a plane that expands towards the user). In some embodiments, the second type of transition is performed regardless of the difference in depth within the individual environment.For example, the transition expands outwardly from the starting position, optionally equally in all directions, regardless of whether parts of the individual environments are closer or farther (e.g., the order in which the object is "enveloped" by the mimicked environment does not depend on the depth of the object within the environment). In some embodiments, the second type of transition includes starting the transition from a location on a surface (e.g., the surface of a table, the surface of a wall, etc.) and moving along the surface (e.g., towards the user). In some embodiments, the second type of transition includes starting the transition from an individual location (e.g., the center of the field of view, etc.) and expanding radially outward along the user's field of view (e.g., expanding in all directions from the individual location in the same manner). In some embodiments, transitioning outwardly from a central reference location provides an initial reference point for the first mimicked environment, which reduces dizziness and / or cognitive dissonance potentially by transitioning from an environment with a short depth (e.g., the user's physical environment) to an environment with a far depth (e.g., the first mimicked environment). In some embodiments, the first and second transition methods have different transition speeds, different cross-fade animations, and / or different orders of transitioning parts of the environment. In some embodiments, the results of the first and second transition methods are the same. For example, the initial and final conditions are the same, but the transition from the initial condition to the final condition is different.
[0159] (For example, by transitioning from a first environment to a second environment in a manner based on the characteristics of the physical environment around the device) The above-described method of transitioning from one environment to another provides a fast and efficient way to transition to different environments in a manner that takes into account the physical environment and / or the characteristics of the starting environment, and optionally reduces the abruptness of the transition, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient (e.g., by enabling the user to interact with the new environment more quickly and reducing the time required to adapt to the change), reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby further reducing power consumption and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0160] In some embodiments, one or more first criteria are met when a portion of the physical environment around the user's perspective in an individual direction relative to the reference orientation of the electronic device has a distance greater than a threshold distance from the user 720 of FIG. 7A to the far wall, such as the physical environment 702, and thus qualifies the three-dimensional environment 704 for a long-distance field transition (e.g., when a portion of the physical environment of the electronic device in front of the electronic device (e.g., in the "front" direction and / or the direction of the center of the field of view of the perspective of the three-dimensional environment displayed by the device) has a space greater than a threshold amount (e.g., greater than 1 foot, 5 feet, 10 feet, 50 feet, 100 feet, etc.), the first criterion is met) (810).
[0161] For example, if the device is located in the center of the room and the wall in front of the device (e.g., the boundary of the physical environment in front of the device) is more than 10 feet away, one or more first criteria are met (e.g., if the wall in front of the device is less than 10 feet away, one or more first criteria are optionally not met). In some embodiments, the boundary of the physical environment is defined by a wall of the physical environment (e.g., a vertical barrier that exceeds a threshold height such as 5 feet, 8 feet, 12 feet, etc., or a vertical barrier that extends the vertical distance from the floor to the ceiling). In some embodiments, objects within the physical environment that are not recognized as vertical walls or are below the threshold height are not considered when determining whether the physical environment is greater than a threshold distance from the electronic device. For example, a sofa, table, desk, counter, etc. in front of the electronic device is optionally not considered when determining whether the physical environment is greater than a threshold distance from the electronic device. In some embodiments, a sofa, table, desk, counter, etc. are considered. In some embodiments, the size and / or shape of the physical environment (e.g., including the size, shape, and / or location of objects within the physical environment) are determined via one or more sensors of the electronic device such as a visible light sensor (e.g., a camera), a depth sensor (e.g., a time-of-flight sensor) (e.g., individually or in combination).
[0162] The above method of displaying the environment via either a first type of transition or a second type of transition (e.g., based on whether the physical environment around the device meets certain criteria) ensures a smooth transition that takes into account the user's environment and reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient (e.g., automatically, without the need for the user to perform additional inputs to select between different types of transitions), thereby reducing power usage further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0163] In some embodiments, one or more second criteria include a criterion that is satisfied when a portion of the physical environment around the user's perspective in an individual direction relative to a reference orientation of the electronic device is less than a threshold distance from the electronic device (812). For example, the physical environment 702 has a distance less than the threshold distance from the user 720 to the back wall in FIG. 7B, and thus qualifies the three-dimensional environment 704 for near-field transition and optionally does not qualify for far-field transition (e.g., a portion of the physical environment of the electronic device in front of the electronic device (e.g., in the "front" direction and / or the direction of the center of the field of view of the perspective of the three-dimensional environment displayed by the device) has a space less than a threshold amount (e.g., less than 1 foot, 5 feet, 10 feet, 50 feet, over 100 feet, etc.), the second criterion is satisfied).
[0164] For example, if a wall in front of the electronic device is less than a threshold distance from the electronic device, the electronic device performs a second type of transition when displaying an individual environment. In some embodiments, the second type of transition is a different type of transition when there is a space exceeding a threshold amount in order to reduce the risk that the transition is uncomfortable for the user and / or to reduce the risk of causing symptoms of dizziness. For example, if the user is presented with an environment with a small amount of depth (e.g., the user's physical environment) and suddenly transitions to an environment with a lot of space (e.g., a simulated environment), this effect may cause visual dissonance. In some embodiments, one or more second criteria include a criterion that is satisfied when the electronic device cannot determine depth information of the physical environment around the electronic device (e.g., optionally, to a threshold accuracy level). For example, when environmental conditions prevent the device's sensors from determining the size and shape of the user's room (e.g., atmospheric interference, electromagnetic interference, etc.), one or more second criteria are satisfied.
[0165] The above method of displaying an environment via either a first type of transition or a second type of transition (e.g., based on whether the physical environment around the device meets certain criteria) ensures a smooth transition that takes into account the user's environment and reduces potential vestibular disorders that could potentially lead to dizziness or intoxication symptoms. This simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient (e.g., automatically, without the need for the user to perform additional inputs to select between different types of transitions), thereby further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.
[0166] In some embodiments, an individual environment includes a representation of a portion of the physical environment around the user's perspective (e.g., at least a portion of the environment presented to and / or displayed to the user is based on the physical environment around the device and / or the user via an actual pass-through via a display generation component (e.g., a transparent or translucent display generation component) or a digital pass-through via a display generation component (e.g., a realistic representation thereof)). Transitioning from displaying an individual environment to displaying a first simulated environment involves replacing the display of a representation of a portion of the physical environment with the display of a representation of a portion of the first simulated environment (e.g., in FIG. 7C, replacing the representation of the corners of the physical environment 702 of the three-dimensional environment 704 with respective portions (e.g., 722-1a and 722-1b) of the virtual environment 722), such as replacing a portion of the representation of the physical environment with a portion of the representation of the simulated environment (e.g., replacing a portion of the representation of the physical environment so that the environment presented to the user is partially the real-world physical environment and partially the virtual simulated environment) (814).
[0167] In some embodiments, the amount of the simulated environment displayed (e.g., the amount of the physical environment not displayed) is based on the device and / or the immersion level of the simulated environment. For example, increasing the immersion level causes more simulated environments to be displayed, replacing and / or obscuring more of the physical environment representation, and decreasing the immersion level causes fewer simulated environments to be displayed, exposing portions of the physical environment that were not previously displayed and / or obscured.
[0168] (For example, by replacing portions of the physical environment representation with portions of the simulated environment) The above-described method of displaying a simulated environment provides a quick and efficient way to transition from displaying only the physical environment to also displaying the simulated environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby improving the battery life of the electronic device by reducing power usage further and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0169] In some embodiments, the input corresponding to a request to display a first simulated environment includes user input (e.g., operating a mechanical input element that communicates with an electronic device such as a dial or button by the user, or detecting and / or receiving an operation of a control element such as a virtual button or virtual dial displayed within an individual environment) such as an affordance or selection of a simulated environment representation while currently displaying a virtual environment such as any of FIGS. 7C-7H (816). In some embodiments, the user input is performed by one or more hands of the user and detected by one or more sensors of the device such as a camera, a motion sensor, a hand-tracking sensor.
[0170] In some embodiments, transitioning from displaying an individual environment to displaying a first emulated environment corresponds to a determination that user input corresponds to a request to transition to the first emulated environment by a first amount (e.g., the user input selects an individual immersion level that is higher or lower than the current immersion level), transitioning from displaying the individual environment to displaying a first portion of the first emulated environment (820), e.g., transitioning from an emulated environment having a first immersion level to an emulated environment having a second immersion level (e.g., transitioning from any of FIGS. 7C-7H to any other of FIGS. 7C-7H) (e.g., displaying an individual amount of the first emulated environment based on the user input), and transitioning from displaying the individual environment to displaying a second portion of the first emulated environment that is larger than the first portion (824) in accordance with a determination that the user input corresponds to a request to transition to the first emulated environment by a second amount that is greater than the first amount, e.g., transitioning from an emulated environment having a first immersion level to an emulated environment having a second immersion level (e.g., transitioning from any of FIGS. 7C-7H to any other of FIGS. 7C-7H) (e.g., if the user input is a request to increase (or decrease) the immersion level by two steps, the amount of the first emulated environment displayed increases (or decreases) by an amount corresponding to the two steps of increased (or decreased) immersion), including (818).
[0171] In some embodiments, the user input is a relative input that increases or decreases the immersion level by a specific amount (e.g., a request to increase the immersion level by one step, two steps, etc.). In some embodiments, the user input is a selection of a specific immersion level (e.g., a request to increase the immersion level to 50%, 75%, etc.). For example, if the user input is a request to increase the immersion level by one step, the amount of the first simulated environment displayed increases by an amount corresponding to one step of increased immersion. In some embodiments, if the user input is a request to increase the immersion level by two steps, the amount of the first simulated environment displayed increases by an amount corresponding to two steps of increased immersion. In some embodiments, increasing or decreasing the amount of immersion includes gradually increasing or decreasing, respectively, the amount of the first simulated environment displayed (as opposed to, for example, more or less rapidly displaying the first simulated environment).
[0172] (For example, by gradually increasing or decreasing based on the amount requested by the user input) The above-described method of increasing or decreasing the amount of the simulated environment displayed provides a quick and efficient way to display more or less of the simulated environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential motion sickness or dizziness symptoms that may potentially lead to potential visual impairment, thereby reducing power consumption further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0173] In some embodiments, the user input includes rotation of a rotatable input element that communicates with the electronic device (e.g., a mechanical dial on the electronic device or on another device that communicates with the electronic device, such as a remote control device, another electronic device, etc.), and the user input corresponding to a request to transition to the first emulation environment by a first amount includes rotating the rotatable input element by a first individual amount (e.g., 1 increment (and / or a first rotation amount), 2 increments (and / or a second larger rotation amount), etc. (e.g., 1 detent, 2 detents, etc.)), and the user input corresponding to a request to transition to the first emulation environment by a second amount includes rotating the rotatable input element by a second individual amount that is greater than the first individual amount, e.g., rotating a mechanical dial attached to or communicating with the electronic device 101 (e.g., rotating the mechanical dial by a greater amount) (826). In some embodiments, rotating in a first direction corresponds to a request to increase the immersion level, and rotating in a second opposite direction corresponds to a request to decrease the immersion level.
[0174] The above-described method of increasing or decreasing the amount of the emulation environment displayed (e.g., by rotating the rotatable element by an amount) provides a quick and efficient way to display more or less of the emulation environment (e.g., by providing the user with a process of increasing or decreasing the current immersion level without the user having to select a specific amount of immersion), which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby improving the battery life of the electronic device by further reducing power usage and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0175] In some embodiments, after transitioning from displaying an individual environment using a first type of transition to displaying a first emulated environment, the display via the display generation component of the first emulated environment is the same as the display via the display generation component of the first emulated environment after transitioning from displaying an individual environment using a second type of transition to displaying the first emulated environment. For example, the virtual environment 722a of FIG. 7H looks the same or similar to FIG. 7H (828) after each of their respective transitions is completed (e.g., the start and end states of the first and second types of transitions are the same or similar).
[0176] For example, when transitioning from a first immersion level to a second immersion level, the system determines whether to perform a first type of transition or a second type of transition. In some embodiments, the first and second types of transitions include different types of animations and different portions of the display and / or environmental transitions at different times, but the final result after the transition is complete is the same regardless of which type of transition is used. In some embodiments, transitioning from a first immersion level to a second immersion level includes transitioning through each intermediate immersion level between the first immersion level and the second immersion level. In some embodiments, the highest immersion level (e.g., the maximum immersion level) looks the same or similar regardless of whether the physical environment meets a first criterion or a second criterion (e.g., regardless of whether a first type of transition or a second type of transition is used). In some embodiments, the intermediate immersion levels (e.g., levels lower than the maximum immersion level and higher than the non-immersion level) are different based on whether the physical environment meets a first criterion or a second criterion. For example, the size and / or shape of the first emulated environment within the three-dimensional environment optionally differ depending on whether the physical environment meets a first criterion or a second criterion, but the size and shape of the first emulated environment at the maximum immersion level are the same regardless of whether the physical environment meets a first criterion or a second criterion.
[0177] (For example, different types of transitions are performed based on the situation, but by reaching the same final state after the transition is completed) The above-described method of displaying a simulated environment (for example, by providing a consistent experience regardless of which type of transition is used to display the simulated environment) provides a quick and efficient way to display the simulated environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.
[0178] In some embodiments, the first type of transition uses the position and distance information of objects within the physical environment 702 of FIGS. 7C, 7E, and 7G to move the virtual environment 722 to a position close to the user 720, such as a long-distance field transition. It proceeds in a manner based on the difference in distance between 1) the electronic device and a first portion of the physical environment surrounding the user's perspective and 2) the electronic device and a second portion of the physical environment surrounding the user's perspective (830) (for example, when determining which part of the individual environment transitions to the first, second, etc. (for example, the order of the transitions), the position of the objects within the user's physical environment is taken into account).
[0179] For example, the first type of transition optionally first replaces a portion of the representation of the physical environment that is farthest from the user (or device) with a portion of the simulated environment, then replaces a portion of the representation of the physical environment that is next farthest from the user (or device), and so on.
[0180] In some embodiments, the second type of transition proceeds in a manner (e.g., not based on) that is independent of the distance and / or location of objects within the physical environment of FIGS. 7D, 7F, and 7H, and is independent of the difference between 1) the distance between the electronic device and a first portion of the physical environment surrounding the user's viewpoint, and 2) the distance between the electronic device and a second portion of the physical environment surrounding the user's viewpoint, such as a virtual environment 722 that extends radially from a central position (e.g., the second type of transition does not consider the position or distance of objects within the room).
[0181] In some embodiments, the second type of transition selects an individual location to initiate the transition, replaces the representation of the physical environment at the starting position with a portion of the mimicked environment, and then extends the replacement process outward from the initial point regardless of whether the object is in front of or behind another object (e.g., the extension is optionally based on the display area rather than what is being displayed). In some embodiments, the starting position is a location within the representation of the physical environment that is the farthest point from the user's viewpoint. For example, if the device is facing a flat wall, the second type of transition optionally starts from the center of the flat wall. In some embodiments, as the second type of transition proceeds, the mimicked environment expands from the starting position. In some embodiments, the mimicked environment expands outward in a circular shape. In some embodiments, the expansion of the mimicked environment has a shape similar to the surface of a sphere (or hemisphere) surrounding the user (e.g., the device and / or the user is at the center point of the sphere (or hemisphere)), such that the mimicked environment expands as if it were along the surface of a sphere (e.g., an imaginary sphere that is not displayed). In some embodiments, as the mimicked environment expands, the boundary of the mimicked environment expands outward (e.g., in the x - y direction) but also forward (e.g., in the z direction), and optionally remains at the same distance from the user (e.g., because the radius of the sphere is constant throughout the sphere).
[0182] The above-described method of displaying a simulated environment (e.g., by transitioning in a manner that takes into account the position and depth information of the environment around the device, or in a manner that does not take into account the position and depth information of the environment around the device) provides a transition style adjusted based on the type of environment in which the device is located, which improves the operability of the electronic device, makes the user-device interface more efficient, and reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication (e.g., without the user having to determine the ideal type of transition and perform additional input to select each type of transition), thereby further reducing power consumption and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0183] In some embodiments, the second type of transition includes a cross-dissolve between the individual environment and the first simulated environment, and the cross-dissolve begins at a portion of the physical environment farthest from the electronic device (834), as shown in FIG. 7D (e.g., the second type of transition begins at a location corresponding to the farthest location within the physical environment).
[0184] In some embodiments, from the starting location, the second type of transition extends the simulated environment outwardly with respect to the display area (e.g., the display screen). For example, from the starting location, the simulated environment expands radially by 1 cm of the display area per second. In some embodiments, cross-dissolving between the individual environment and the first simulated environment includes fading out the individual environment while simultaneously fading in in the first simulated environment. In some embodiments, the individual environment and the first simulated environment are displayed simultaneously during part of the transition (e.g., when the individual environment is fading out and when the first simulated environment is fading in). In some embodiments, cross-dissolving between the individual environment and the first simulated environment includes continuously replacing each part of the individual environment with each part of the first simulated environment until all respective parts of the individual environment are completely replaced by the first simulated environment.
[0185] The above-described method of displaying the simulated environment (e.g., by cross-fading outwardly from the starting location from the individual environment to the first simulated environment) provides a smooth and efficient transition, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby improving the battery life of the electronic device by further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0186] In some embodiments, transitioning from displaying an individual environment using a first type of transition to displaying a first mimicked environment involves, during a first part of the transition (838), according to a determination that an individual part of the physical environment around the user's perspective is greater than a first distance from the user's perspective, replacing, via a display generation component, the display of the representation of the individual part of the physical environment with the display of at least a first part of the first mimicked environment (840), e.g., determining that a part of the three-dimensional environment 704 corresponds to a part of the physical environment 702 that is farther than an individual distance associated with the immersion level of 3 in FIG. 7E, and thus replacing the determined part with the virtual environment 722a (e.g., if the first part of the representation of the physical environment corresponds to a location that is farther from the user (or device) than a second part, the first part of the representation of the physical environment is replaced (e.g., transitioned) with a part of the first mimicked environment before the second part is replaced (e.g., transitioned)), and according to a determination that an individual part of the physical environment around the user's perspective is less than the first distance from the user's perspective, refraining from replacing, via the display generation component, the display of the representation of the individual part of the physical environment with the display of at least a first part of the first mimicked environment (842), e.g., determining that a part of the three-dimensional environment 704 corresponds to a part of the physical environment 702 that is closer than an individual distance associated with the immersion level of 3 in FIG. 7E, and thus maintaining those parts as the representation of the physical environment (e.g., if the first part of the representation of the physical environment is not far from the second part, during the first part of the transition (e.g., until after the second part has transitioned), not replacing (e.g., transitioning) the first part with a part of the first mimicked environment), and includes (836).
[0187] In some embodiments, during a second portion of the transition (e.g., after a first portion of the transition) (844), in accordance with a determination that an individual portion of the physical environment surrounding the user's perspective is greater than a second distance different from a first distance from the user's perspective, the display of the representation of the individual portion of the physical environment is replaced via a display generation component with the display of at least a second portion of the first mimicked environment (846), e.g., determining that a portion of the three-dimensional environment 704 corresponds to a portion of the physical environment 702 that is farther than an individual distance associated with the immersion level of 6 in FIG. 7G, and thus replacing the determined portion with the virtual environment 722a (e.g., during the second portion of the transition, a second portion of the representation of the physical environment closer than the first portion (e.g., the next farthest portion after the first portion) is replaced with (e.g., transitioned to) a portion of the first mimicked environment before the second portion is replaced (e.g., transitioned), and ceasing to replace the display of the representation of the individual portion of the physical environment with the display of at least a second portion of the first mimicked environment via the display generation component in accordance with a determination that an individual portion of the physical environment surrounding the user's perspective is less than the second distance from the user's perspective (848), e.g., determining that a portion of the three-dimensional environment 704 corresponds to a portion of the physical environment 702 that is closer than an individual distance associated with the immersion level of 6 in FIG. 7H, and thus maintaining those portions as the representation of the physical environment (e.g., not replacing the second portion of the representation if the second portion of the representation of the physical environment is not the next farthest portion after the first portion, (e.g., until the second portion becomes the next farthest portion that has not yet transitioned)).
[0188] In some embodiments, the boundary between the physical environment and the first simulated environment has a shape defined by the surface of a sphere around the user's perspective (e.g., having the device and / or the user as a center point). In some embodiments, as the transition progresses, the sphere moves closer to the electronic device and / or the user and / or the size of the sphere decreases. In some embodiments, a portion of the surface of the sphere that intersects the physical environment defines the boundary between the physical environment and the simulated environment (e.g., the intersection leads to the simulated environment and before the intersection is the physical environment). In some embodiments, as the sphere approaches the electronic device and / or the user and / or as the size of the sphere decreases, more of the surface of the sphere intersects the physical environment, and thus the simulated environment appears to encompass more of the three-dimensional environment and move closer to the user's perspective (e.g., starting from the farthest location within the physical environment).
[0189] The above-described method of transitioning from displaying an individual environment to displaying a simulated environment (e.g., by transitioning a portion of a farther individual environment before transitioning a portion of a nearer individual environment) provides a smooth and efficient transition, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential disorientation disorders that can potentially lead to symptoms of dizziness or intoxication, thereby improving the battery life of the electronic device by further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0190] In some embodiments, during the first part (852) of the transition, according to a determination that an individual part of the physical environment around the user's perspective is greater than a first distance from the user's perspective (854), and according to a determination that a second individual part of the physical environment around the user's perspective is less than the first distance from the user's perspective, via a display generation component, display a representation of the second individual part of the physical environment between the first part of the first simulated environment and the user's perspective (856), for example, determine that a part of the three-dimensional environment 704 corresponds to a part of the physical environment 702 that is closer than an individual distance associated with the immersion level of 3 in FIG. 7E, and maintain those parts as a representation of the physical environment (for example, during the first part of the transition, if an object (where a representation exists in an individual environment) within the physical environment of the device is at a closer distance than the simulated environment, the representation of the object is displayed in front of the simulated environment (for example, architectural objects, furniture, etc.)), and according to a determination that a second individual part of the physical environment around the user's perspective is greater than the first distance from the user's perspective, stop displaying the representation of the second individual part of the physical environment via the display generation component (858), for example, determine that a part of the three-dimensional environment 704 corresponds to a part of the physical environment 702 that is farther than an individual distance associated with the immersion level of 3 in FIG. 7E, and replace those parts with individual parts of the virtual environment 722 (for example, objects farther than the simulated environment are no longer displayed), and include (850). For example, the object is overlaid, obscured, and / or replaced by the simulated environment.
[0191] In some embodiments, if the representation of an object is located along the line of sight of a part of the simulated environment, the representation of the object is displayed overlaid on the part of the simulated environment. Thus, the objects of the physical environment are optionally displayed overlaid on (and / or obscuring) the parts of the simulated environment and / or the ambiguous parts of the simulated environment if each object is closer to the user / device than the simulated environment.
[0192] Display an individual environment simultaneously with the representation of the physical environment (e.g., when an individual object is located between the user / device and the simulated environment, display the representation of the object and optionally obscure portions of the individual environment). The method described above provides a smooth and efficient transition, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby reducing power consumption further and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.
[0193] In some embodiments, the individual environment includes a representation of the physical environment surrounding the user's perspective, and the first simulated environment includes a first simulated light source (860) (e.g., the individual environment simultaneously includes representations of the physical environment and the simulated environment, and the simulated environment includes at least one light source). For example, the simulated environment is an outdoor environment having the sun as a light source, or the simulated environment is a room having a lamp as a light source, or an application user interface (e.g., a movie screen) that emits simulated light.
[0194] In some embodiments, while simultaneously displaying at least a portion of the representation of the physical environment surrounding the user's perspective and at least a portion of the first simulated environment (e.g., including the first simulated light source) via a display generation component, the electronic device projects an illumination effect from the virtual environment 722a onto the real-world environment portion of the three-dimensional environment 704 of FIG. 7E (e.g., when a light source in the simulated environment projects an illumination effect onto the representation of the physical environment), display a portion of the representation of the physical environment with an illumination effect based on the first simulated light source (862). For example, light from the sun in the simulated environment is optionally displayed as if it exits the simulated environment, enters the physical environment, and illuminates a portion of the representation of the physical environment. In some embodiments, the simulated light source optionally displays a simulated shadow within the representation of the physical environment.
[0195] (For example, by displaying an illumination effect on a portion of the physical environment resulting from a light source within the simulated environment) The above-described method of displaying an illumination effect from a simulated environment within a representation of a physical environment (for example, by blending aspects of the simulated environment with the physical environment) increases the immersion effect of the simulated environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby reducing power consumption further and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.
[0196] In some embodiments, the illumination effect is based on the geometric shape of the physical environment around the user's viewpoint, such as the shape and position of rooms and objects within the physical environment 702 of FIG. 7E (864) (for example, the illumination effect from a simulated light source within the simulated environment is displayed as if it were interacting with objects within the physical environment according to physics). For example, an object within the line of sight of the light source experiences the illumination effect, while an object obscured (for example, by a virtual object or a physical object) optionally does not experience the illumination effect (for example, and optionally has a shadow projected thereon by an object within the line of sight of the light source). In some embodiments, some objects receive a scattered, reflected, or refracted illumination effect.
[0197] (For example, by displaying an illumination effect on a portion of the physical environment based on the physical characteristics of light from an illumination source) The above-described method of displaying an illumination effect from a simulated environment in the representation of the physical environment increases the immersion effect of the simulated environment (e.g., by realistically projecting light from the illumination source onto the physical environment), provides information and / or instructions about what the device knows about the physical environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential sense-of-orientation impairments that can potentially lead to symptoms of dizziness or intoxication, thereby reducing power usage further, improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0198] In some embodiments, the illumination effect is based on the architecture of the physical environment around the user's viewpoint, such as the shape and position of rooms and objects within the physical environment 702 of FIG. 7E (866) (e.g., the illumination effect from a simulated illumination source within the simulated environment is displayed as if it interacts with the boundaries of the physical environment according to physics). For example, the illumination source optionally projects light onto walls and through windows.
[0199] (For example, by displaying an illumination effect on a portion of the physical environment based on the physical characteristics of light from an illumination source) The above-described method of displaying an illumination effect from a simulated environment in the representation of the physical environment increases the immersion effect of the simulated environment (e.g., by realistically projecting light from the illumination source onto the physical environment), provides information and / or instructions about what the device knows about the physical environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential sense-of-orientation impairments that can potentially lead to symptoms of dizziness or intoxication, thereby reducing power usage further, improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0200] In some embodiments, the individual environment includes a representation of the physical environment surrounding the user's perspective, and the first mimetic environment includes a first volumetric effect such as the spatial effect of FIG. 9B (e.g., an effect that changes the visual characteristics of the space within the environment not occupied by an object) (868). For example, an effect that changes the appearance and feel of the air within the environment. In some embodiments, the volumetric effects include lighting effects, fog effects, smoke effects, and the like.
[0201] In some embodiments, while simultaneously displaying at least a portion of the representation of the physical environment surrounding the user's perspective and at least a portion of the first mimetic environment (e.g., including the first volumetric effect) via a display generation component, the electronic device displays a spatial effect on the real-world portion of the three-dimensional environment 704 based on the virtual environment 722 of FIG. 7E within the volume of a portion of the representation of the physical environment (e.g., displays a volumetric effect within the representation of the physical environment based on the volumetric effect within the mimetic environment), such as a second volumetric effect based on the first volumetric effect (870). For example, if the mimetic environment includes a fog effect, a fog effect is optionally displayed within the representation of the physical environment such that the fog from the mimetic environment appears to drift into the physical environment.
[0202] The above-described method of displaying a volumetric effect from the mimetic environment in the representation of the physical environment (e.g., by displaying a volumetric effect on a portion of the physical environment based on the volumetric effect shown in the mimetic environment) increases the immersion effect of the mimetic environment (e.g., by causing the effect from the mimetic environment to flow out into the physical environment), which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby improving the battery life of the electronic device by further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0203] In some embodiments, the first volumetric effect and the second volumetric effect include an illumination effect (874) (e.g., the volumetric effect includes changing the illumination and / or introducing an illumination effect). For example, a fog volumetric effect includes changing the illumination of virtual particles displayed within a physical environment to mimic the scattering / diffraction effect of fog.
[0204] (For example, by displaying an illumination effect on a portion of the physical environment based on the illumination effect within the mimicked environment) The above method of displaying an illumination effect from a mimicked environment within a representation of a physical environment (e.g., by realistically projecting light from an illumination source onto the physical environment) increases the immersion effect of the mimicked environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby reducing power consumption further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0205] In some embodiments, the first volumetric effect and the second volumetric effect include a particle effect (874) (e.g., displaying visual effects (e.g., abstract visual effects) such as fire, smoke, dust, etc. in a physical environment). In some embodiments, the particle effect includes one or more mimicked particles in the atmosphere of a physical environment that emit one or more visual effects.
[0206] (For example, by displaying a particle effect on a portion of the physical environment based on a particle effect within a simulated environment) The above-described method of displaying a particle effect from a simulated environment within a representation of a physical environment (for example, by realistically displaying a particle effect within the physical environment that is an optional extension of the particle effect displayed within the simulated environment) increases the immersion effect of the simulated environment, which improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby reducing power consumption further and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.
[0207] In some embodiments, while simultaneously displaying at least a portion of an individual environment and at least a portion of a first simulated environment via a display generation component (876), in accordance with a determination that a boundary between a portion of the individual environment and a portion of the first simulated environment intersects a representation of an individual object within the individual environment, the electronic device removes the representation of the individual object from the display via the display generation component (878) (for example, when the boundary of the simulated environment intersects the representation of a physical object within the representation of the physical environment (e.g., as opposed to the object being completely in front of or completely behind the simulated environment), completely removing the display of the object). In some embodiments, instead of removing the display of the object, a portion of the object that is not behind the boundary of the simulated environment is displayed (e.g., only a portion of the object is displayed and the portion of the object that is behind the simulated environment is hidden).
[0208] Simultaneously display a representation of a virtual environment and a physical environment (e.g., by removing objects in the physical environment that would otherwise be partially within the physical environment and partially within the virtual environment), the method described above increases the immersion effect of the virtual environment (e.g., by avoiding scenarios where the fusion between the virtual environment and the physical environment is unnatural), improves the operability of the electronic device, makes the user-device interface more efficient, reduces potential vestibular disorders that can potentially lead to symptoms of dizziness or intoxication, thereby improving the battery life of the electronic device by further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0209] Figures 9A - 9H show examples of changing (or not changing) the immersion level of a three-dimensional environment according to some embodiments.
[0210] Figure 9A shows an electronic device 101 that displays a three-dimensional environment 904 on a user interface via a display generation component (e.g., the display generation component 120 of FIG. 1). As described above with reference to FIGS. 1 - 6, the electronic device 101 optionally includes a display generation component (e.g., a touch screen) and a plurality of image sensors (e.g., the image sensors 314 of FIG. 3). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the user interface shown below can also be implemented on a head-mounted display that includes a display generation component for displaying the user interface to the user and sensors for detecting the physical environment and / or the movement of the user's hand (e.g., external sensors facing outward from the user), and / or the user's line of sight (e.g., internal sensors facing inward towards the user's face).
[0211] As shown in FIG. 9A, device 101 captures one or more images of the real-world environment 902 (e.g., operating environment 100) around device 101, including one or more objects in the real-world environment 902 around device 101. In some embodiments, device 101 displays a representation of the real-world environment within a three-dimensional environment 904. For example, the three-dimensional environment 904 includes a representation of a portion of a desk 910a, a representation of a coffee table 914a, and a representation of a side table 912a. As shown in the overhead view 918 of the real-world environment 902 in FIG. 9A, the corner table 908b, desk 910b, side table 912b, and coffee table 914b are real objects in the real-world environment 902 captured by one or more sensors of device 101. In FIG. 9A, the three-dimensional environment 904 is currently at an immersion level of 3 (as indicated, for example, by immersion level indicator 916), and as a result, the virtual environment 922a encompasses a portion of the three-dimensional environment 904 and obscures the view of the corner table 908b and the back portion of the desk 910b (in a manner similar to the three-dimensional environment 704 described above with reference to FIGS. 7A-7H, such as FIG. 7E).
[0212] In some embodiments, when the three-dimensional environment 904 is at an individual immersion level (e.g., a non-immersed level), device 101 generates one or more audio effects associated with the virtual environment 922a (as indicated, for example, by speaker 924). In some embodiments, the audio effects include ambient acoustic effects audible in the virtual environment 922a. For example, a virtual environment of a peaceful pasture may include the sounds of frogs, birds, running water, etc., that are generated by device 101 as if they were audible while the user is within the physical environment 902, although they do not exist in the physical environment 902 by desire (e.g., the surrounding physical environment does not include these sounds). In some embodiments, providing the audio effects increases the immersion effect of the virtual environment 922a and thus gives the user the experience of being within the virtual environment 922a.
[0213] In some embodiments, the three-dimensional environment 904 is displayed at an intermediate immersion level, as shown in FIG. 9A, in response to user input requesting the display of a virtual environment (e.g., virtual environment 922a, etc.). For example, the device 101 can display one or more selectable representations of a virtual environment (e.g., a buffet of virtual environments) that are selectable to display the virtual environment selected in the immersive experience as if the user were physically located within the selected virtual environment. In some embodiments, in response to user input selecting an individual virtual environment, the device 101 displays the selected virtual environment at a default immersion level that is below the maximum immersion level (e.g., the highest allowable immersion level optionally limited by system settings or user-defined settings) and higher than the minimum immersion level (e.g., the immersion level just above no immersion). In some embodiments, displaying the selected virtual environment at a default intermediate immersion level gently guides the user into an immersive virtual environment and allows the user to increase or decrease the immersion level as desired. In some embodiments, the user can increase or decrease the immersion level by rotating a mechanical dial in one direction or the other, or by selecting an affordance to increase or decrease the immersion level. In some embodiments, the user can select an individual immersion level by interacting with the immersion level indicator 916. For example, in response to the user performing a selection input on an element of the immersion level indicator 916 at the maximum immersion level (e.g., the rightmost square), the three-dimensional environment 904 displays the virtual environment 922a at the maximum immersion level, but in response to a selection input directed towards the lowest immersion level element of the immersion level indicator 916 (e.g., the leftmost square), the virtual environment 922a is displayed at the lowest immersion level. In some embodiments, the transition from one immersion level to another is described above with reference to method 800.
[0214] Figure 9B shows an embodiment in which an immersive spatial effect (e.g., as contrasted with an immersive virtual environment) is displayed at an immersion level of 3. In some embodiments, a spatial effect refers to artificially (e.g., virtually) modifying the physical atmosphere within the three-dimensional environment 904 or the visual characteristics of one or more objects without optionally replacing or removing one or more objects within the three-dimensional environment 904. For example, a spatial effect can include displaying the three-dimensional environment 904 as if it were illuminated by sunset, displaying the three-dimensional environment 904 as if it were foggy, displaying the three-dimensional environment 904 as if it were dimly lit, and the like. As shown in the overhead view 918, displaying a spatial effect does not optionally include displaying a virtual environment in place of a portion of the real-world environment portion of the three-dimensional environment (e.g., it is understood that both an immersive spatial effect and an immersive virtual environment can be displayed simultaneously, at the same immersion level or different immersion levels). In Figure 9B, the spatial effect includes displaying virtual light entering the physical environment from the window 915. In some embodiments, the virtual light entering from the window 915 has an illumination effect that is different from the amount of real-world light entering from the window 915 within the real-world environment 902 (e.g., the virtual light is virtual and exists only within the three-dimensional environment 904). In some embodiments, the virtual light causes an illumination effect to be displayed within the three-dimensional environment 904 according to particle physics. For example, the virtual illumination optionally causes a glare to occur on the picture frame 906 on the back wall of the room, projects a shadow 926 onto the table 910a, and projects a shadow onto the coffee table 914a, as shown in Figure 9B (e.g., one or more of these optionally do not exist within the physical environment 902 or exist within the physical environment 902 with different characteristics such as lower or higher brightness, smaller or larger size). In some embodiments, the virtual light optionally increases or decreases the ambient light within the three-dimensional environment 904 compared to the amount of actual light existing within the physical environment 902.
[0215] FIG. 9C shows a three - dimensional environment 904 in response to an input that increases the immersion level of the virtual environment 922 from a level of 3 (e.g., as in FIG. 9A) to a level of 6 (e.g., as indicated by the immersion level indicator 916). In some embodiments, increasing the immersion level expands the virtual environment 922a to encompass more of the three - dimensional environment 904 and, thus, replaces the representation of the real - world environment 902, as shown in FIG. 9C (e.g., in a manner similar to that described above with respect to method 800). In some embodiments, increasing the immersion level optionally includes increasing one or more audio effects generated by the device 101 (e.g., as indicated by a speaker 924 that provides a louder volume than in FIG. 9A). In some embodiments, increasing one or more audio effects includes increasing the volume of an audio effect provided at a lower immersion level, providing one or more new audio effects, and / or increasing the complexity and / or depth of the audio effects. For example, at a lower immersion level, the device 101 provides an audio effect of a small stream (e.g., associated with the virtual environment 922a), at a higher immersion level, the device 101 adds an audio effect of a bird's chirping (e.g., associated with the virtual environment 922a) to the babbling brook audio effect, and at an even higher immersion level, the device 101 adds an audio effect of a frog's croaking (e.g., associated with the virtual environment 922a) to the audio effects of the small stream and the bird's chirping. In some embodiments, the device 101 increases or does not increase (e.g., maintains or decreases) the volume of the generated audio effect when adding the above - described audio effects to the audio it generates.
[0216] FIG. 9D shows a three-dimensional environment 904 in response to an input that increases the immersion level of the immersive spatial effect from an immersion level of 3 (e.g., as in FIG. 9B) to an immersion level of 6 (e.g., as indicated by the immersion level indicator 916). In some embodiments, increasing the immersion level increases the magnitude of the spatial effect. In some embodiments, increasing the magnitude of the spatial effect includes increasing the amount of modification and / or deformation performed on the real-world environment 902 and / or the three-dimensional environment 904. For example, in FIG. 9D, the virtual light entering through the window 915 (e.g., a virtual modification of the real-world environment 902 as described above) increases in magnitude (e.g., increases in brightness). In some embodiments, increasing the virtual light incident from the window 915 results in more ambient light within the three-dimensional environment 904 and increases the lighting effect. For example, in FIG. 9D, the glare on the picture frame 906 increases as a result of the increased virtual light from the window 915, and the shadow 926 is darker (e.g., due to the increased contrast between the ambient lighting of the three-dimensional environment 904 and the shadow projected by the desk 910a). In some embodiments, increasing the immersion level of the immersive spatial effect includes increasing one or more audio effects associated with the spatial effect in a manner similar to increasing the audio effects described with respect to FIG. 9C for the virtual environment 922a.
[0217] Figures 9E - 9F illustrate an embodiment where, for example, as a result of device 101 being oriented in different directions without a change in the orientation of user 920's body (e.g., torso), the perspective and / or view of device 101 changes. For example, in Figures 9E - 9F, the user moves their hand to turn device 101 to the left (as opposed to forward as in Figures 9A - 9D, for example), but does not rotate their shoulder and / or torso to face left. As another example in the case of head - mounted device 101, the user turns their head to the left without rotating their shoulder and / or torso in conjunction with the rotation of their head. As described above, rotating device 101 causes device 101 to face different directions and uses one or more sensors of device 101 to capture different portions of the real - world environment 902. In some embodiments, in response to detecting that device 101 has changed its orientation such that it faces different portions of the real - world environment 902, device 101 updates the three - dimensional environment to rotate in accordance with the rotation of device 101. For example, in Figures 9E - 9F, since device 101 is currently facing left, the three - dimensional environment 904 displays a representation of the real - world environment 902 that is to the left of user 920 and optionally does not display a portion of the real - world environment 902 that is in front of user 920 as if the user had turned their head to the left to view the portion of the real - world environment 902 that is to the user's left (e.g., the amount that is no longer displayed is optionally based on the field of view of the three - dimensional environment 904 and the amount of rotation that device 101 experiences).
[0218] In FIGS. 9E - 9F, the immersion level is increasing to the maximum immersion level. In FIG. 9E, increasing the immersion level to the maximum optionally includes increasing the amount of the three - dimensional environment 904 occupied by the virtual environment 922a to the maximum amount, which optionally, as reflected in the bird's - eye view 918 of FIG. 9E, optionally replaces all of the three - dimensional environment 904 except for a predetermined radius (e.g., for safety purposes, a radius such as 6 inches, 1 foot, 2 feet, 5 feet, etc.) around the user 920 with the virtual environment 922a, and optionally increasing one or more audio effects to the maximum amount in one or more of the manners described with reference to FIG. 9C.
[0219] In FIG. 9E, since the device 101 is oriented to face different directions, the view of the three-dimensional environment 904 displayed by the device 101 is rotated to display a new perspective view of the three-dimensional environment 904 as a result of the new direction. For example, since the device 101 is currently facing left, the view of the three-dimensional environment 904 displayed by the device 101 includes a view of the left side of the real-world environment 902, including a representation of the side table 923a. In some embodiments, since the body (e.g., torso) of the user 920 did not rotate to face left, the virtual environment 922a does not rotate in response to the change in the orientation of the device 101 and remains oriented in the direction (e.g., forward) in which the body of the user 920 is facing (e.g., remains centered) (e.g., the virtual environment 922a remains at the same location within the three-dimensional environment 904 that was being displayed prior to the change in the orientation of the device 101). In some embodiments, as shown in FIG. 9E, by rotating the device 101 to face left, optionally, the boundary between the virtual environment 922a and the real-world environment 902 is exposed on the device 101. In some embodiments, in response to detecting that the device 101 has rotated by a threshold amount (e.g., rotated by 30 degrees, 45 degrees, 60 degrees, 90 degrees, 120 degrees, etc.), one or more audio effects are reduced (or removed), as indicated by the speaker indicator 924 in FIG. 9E. In some embodiments, one or more audio effects are reduced in proportion to the magnitude of the change in the orientation of the device 101 as the device 101 rotates within the real-world environment 902 and are completely eliminated when the device 110 rotates by a threshold amount. In some embodiments, reducing one or more audio effects includes reducing the volume of one or more audio effects, reducing the complexity of one or more audio effects, pausing one or more audio tracks, or completely pausing all audio effects associated with the virtual environment 922a (e.g., in one or more ways opposite to those described with reference to FIG. 9C).In some embodiments, not rotating or moving the virtual environment 922a within the three-dimensional environment 904 and / or reducing one or more audio effects such that it is oriented (e.g., centered) with the view of the three-dimensional environment 904 presented by the device 101 reduces the amount of immersion experienced by the user and thus enables the user to interact with their physical environment. For example, the user may rotate their head to interact with objects or people within their environment (e.g., thus rotating the device 101 to face in a new direction if the device 101 is a head-mounted device), and by not rotating the virtual environment 922a to be oriented with the view of the three-dimensional environment 904 presented by the device 101 and / or by reducing one or more audio effects, the user can interact with objects or people within their environment without performing an explicit input to reduce the level of immersion in the virtual environment 922a. In some embodiments, subsequently rotating the device 101 by an amount less than a threshold rotation amount optionally causes one or more audio effects to return to their previous levels (e.g., levels associated with a maximum immersion level), for example, gradually, depending on the difference between the current orientation of the device 101 and the location of the virtual environment 922a within the three-dimensional environment 904.
[0220] In FIG. 9F, device 101 is reoriented as described with reference to FIG. 9E. In FIG. 9F, since device 101 is oriented to face a different direction, the view of the three-dimensional environment 904 displayed by device 101 is rotated to display a new perspective as a result of the new direction. In FIG. 9F (e.g., as described with reference to FIGS. 9B and 9D), device 101 is displaying an immersive spatial effect, and since device 101 is displaying an immersive spatial effect, the user cannot see the boundary between the representation of the real-world environment and the virtual environment. As shown in FIG. 9F, side table 923a projects a virtual shadow 928 as a result of virtual light entering through window 915 (e.g., as described above in FIGS. 9B and 9D). Additionally or alternatively, device 101 optionally does not reduce one or more audio effects. In some embodiments, the immersive spatial effect is applied to the entire physical environment of the user, and thus rotating device 101 so that it is oriented towards different parts of the real-world environment does not reduce the amount and / or magnitude of the spatial effect provided by device 101, and thus rotating device 101 enables the user to see the effect of the spatial effect on different areas of their physical environment. In some embodiments, even when providing an immersive spatial effect (e.g., as in the case of displaying a virtual environment), the real-world environment is not made invisible, so even when maintaining the spatial effect, the user is not prevented from interacting with objects or people within the user's environment. For example, as a result of displaying an immersive spatial effect, objects or people within the user's environment are not removed from the display.
[0221] Figures 9G - 9H illustrate an embodiment where, for example, as a result of the rotation of the body of user 920 (e.g., both the device 101 and the shoulders and / or torso of user 920 rotate within the real - world environment 902), the viewpoint and / or perspective view of device 101 changes as compared to FIGS. 9C and 9D. For example, in FIGS. 9G - 9H, user 920 rotates their body to face left while continuing to hold device 101 outwardly, thus causing device 101 to also face left as shown in the overhead view 918. In some embodiments, device 101 can determine the orientation of the body of user 920 (e.g., torso and / or shoulders) via one or more sensors such as a visible - light sensor (e.g., a camera) and / or a depth sensor that faces the user.
[0222] In FIG. 9G, in response to detecting that the device 101 and the user 920's body have rotated to face left, the device 101 moves (e.g., rotates) the virtual environment 922a within the three-dimensional environment 904 so that it is oriented (e.g., centered) using the orientation of the device 101 and / or the user 920's body. In some embodiments, since the device 101 is oriented to face the same direction as the user's body, the virtual environment 922a is placed at the center of the view of the three-dimensional environment 904 displayed by the device 101 such that the entire view of the three-dimensional environment 904 remains occupied by the virtual environment 922a (e.g., excluding the “notch” area around the user 920 as described above). In some embodiments, the device 101 additionally or alternatively maintains one or more audio effects associated with the virtual environment 922a. Thus, as described above, if the device 101 is rotated without the user's body rotating while the virtual environment is being displayed at an individual immersion level, even if the device 101 is no longer oriented in the direction towards the virtual environment (e.g., the user can no longer see the virtual environment and / or the device 101 is no longer displaying the virtual environment due to displaying a different part of the three-dimensional environment 904), the virtual environment is maintained at its absolute position within the three-dimensional environment 904. However, if the device 101 rotates with the user's body in the same or a similar manner (e.g., by the same or a similar amount, in the same or a similar direction, etc.) as the device 101 and / or the user's body, the virtual environment is displayed at a new location within the three-dimensional environment 904 and rotates (e.g., maintains the same location relative to the orientation of the user's body and optionally rotates using an acceleration / deceleration mechanism to avoid a jerky animation) so that the virtual environment is oriented towards the user (e.g., centered).If the user's body rotates without the device 101 rotating (e.g., due to the device 101 not being oriented in a new direction), even if the orientation of the view of the three-dimensional environment 904 does not change, the virtual environment is optionally moved within the three-dimensional environment 904 in accordance with the rotation of the user's body, which can optionally rotate the virtual environment out of view (e.g., based on how much the user's body has rotated and the field of view of the three-dimensional environment 904). In some embodiments, if the user's body rotates without the device 101 rotating, the virtual environment does not rotate in accordance with the rotation of the user's body (e.g., it is understood that the virtual environment optionally rotates only when both the user's body and the device rotate, but they do not need to rotate simultaneously).
[0223] In FIG. 9H, in response to detecting that the device 101 and the user 920's body have rotated to face left (e.g., in a manner similar to FIG. 9G), the view of the three-dimensional environment 904 displayed by the device 101 is rotated to display a new perspective as a result of the new orientation of the device 101, and the device 101 continues to display the immersive spatial effect in a manner similar to FIG. 9F. As shown in FIG. 9H, the side table 923a projects a virtual shadow 928 as a result of virtual light entering through the window 915 (e.g., as described above in FIGS. 9B, 9D, and 9F). Additionally or alternatively, the device 101 optionally does not reduce one or more audio effects. Thus, as described above with reference to FIGS. 9F and 9H, if the device 101 is providing the immersive spatial effect at an individual immersion level, in response to the rotation of the device 101 to face a new direction, regardless of whether the user's body rotates with the device 101, the immersive spatial effect continues to be provided without increasing or decreasing the amount of immersion (e.g., without increasing or decreasing the audio effect, without increasing or decreasing the visual effect, etc.). For example, the three-dimensional environment 904 of FIG. 9H is the same and / or similar to the three-dimensional environment 904 of FIG. 9F, and / or the audio effect provided in FIG. 9H is the same and / or similar to the audio effect provided in FIG. 9F.
[0224] Thus, as described above, device 101 can provide an immersive environment (e.g., as described with respect to FIGS. 9A, 9C, 9E, and 9G) and / or an immersive atmosphere effect (e.g., as described above with respect to FIGS. 9B, 9D, 9F, and 9H). In some embodiments, device 101 can provide both an immersive environment and an immersive atmosphere effect simultaneously (e.g., optionally, affecting only the visual characteristics of the real-world environment portion of the three-dimensional environment 904, or optionally, affecting the visual characteristics of both the real-world environment portion and the simulated environment portion of the three-dimensional environment 904). In such embodiments, the immersive environment exhibits the behavior described above with respect to FIGS. 9A, 9C, 9E, and 9G, and the atmosphere effect exhibits the behavior described above with respect to FIGS. 9B, 9D, 9F, and 9H).
[0225] FIGS. 10A - 10P are flowcharts showing a method 1000 for changing the immersion level of a three-dimensional environment according to some embodiments. In some embodiments, method 1000 is executed in a computer system (e.g., the computer system 101 of FIG. 1 such as a tablet, smartphone, wearable computer, or head-mounted device) that includes a display generation component (e.g., the display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera facing downward with the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras), or a camera facing forward from the user's head). In some embodiments, method 1000 is stored in a non-transitory computer-readable storage medium and is controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., the control unit 110 of FIG. 1A). Some operations of method 1000 can optionally be combined and / or the order of some operations can optionally be changed.
[0226] In method 1000, in some embodiments, an electronic device (e.g., computer system 101 of FIG. 1) that communicates with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer) while displaying a three-dimensional environment without an individual spatial effect (e.g., a representation such as a physical environment, a simulated environment, computer-generated reality, extended reality, etc.) via the display generation component detects (1002) a first input corresponding to a request to initiate an individual spatial effect, such as a selection of an affordance of a representation of an individual spatial effect, via one or more input devices while the three-dimensional environment 904 is not displaying a spatial effect, as shown in FIG. 7A or FIG. 7B (e.g., detecting a movement of one or more hands of the user, a movement of one or more eyes of the user and / or a change in line of sight, and / or any other input corresponding to a request to initiate a spatial effect via a hand tracking device, an eye tracking device, or any other input device).
[0227] In some embodiments, the display generation component is a display integrated with an electronic device (optionally a touch screen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface so that the user interface is visible to one or more users. In some embodiments, one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touch screen, a mouse (e.g., external), a track pad (optionally integrated or external), a touch pad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor and / or a motion sensor (e.g., a hand tracking sensor, a hand motion sensor), etc.
[0228] In some embodiments, the three-dimensional environment is an augmented reality environment or a mixed reality environment that optionally includes virtual objects and / or representations of real-world objects in the physical world surrounding the electronic device. In some embodiments, the individual environment is at least based on the physical environment surrounding the electronic device. For example, the electronic device can capture visual information about the environment surrounding the user (e.g., objects in the environment, the size and shape of the environment, etc.) and display at least a portion of the physical environment surrounding the user to the user, optionally making it appear as if the user is still located within the physical environment. In some embodiments, the three-dimensional environment is actively displayed to the user via a display generation component. In some embodiments, the three-dimensional environment is passively presented to the user via a partially transparent or translucent display through which the user can view at least a portion of the physical environment. In some embodiments, the spatial effect is an ambient transformation that modifies one or more visual characteristics of the three-dimensional environment so that it appears as if the three-dimensional environment is located in different times, places, and / or conditions (e.g., morning lighting instead of afternoon lighting, sunny instead of cloudy, fog, etc.). For example, the spatial effect can include changes in the warmth of the environment (e.g., the warmth of the light source), changes in the lighting of the environment (e.g., location, type, and / or brightness), changes in the shadows within the environment (e.g., mimicking the addition, removal, or change in position of a light source), etc. In some embodiments, the spatial effect optionally does not add or remove objects from the three-dimensional environment. In some embodiments, the electronic device can contextually recognize a window and modify what is visible through the window according to the spatial effect. For example, if the spatial effect includes displaying the three-dimensional environment as if it were in a tropical rainforest, the atmosphere of the three-dimensional environment is modified to appear dimly lit, wet, etc., and the window is modified so that a scene of the tropical rainforest is visible through the window.In some embodiments, the spatial effect includes presenting a simulated environment within or simultaneously with a three-dimensional environment (e.g., in a manner similar to that described above with respect to method 800).
[0229] In some embodiments, the request includes selection of affordances to increase the immersion level of the device and / or the individual environment, and / or operation of (e.g., mechanical) control elements (similar to those described with reference to method 800), and / or selection of options associated with individual spatial effects (e.g., from a plurality of available spatial effects). In some embodiments, the request includes a predetermined gesture recognized as a request to initiate a spatial effect. In some embodiments, the request includes voice input requesting display of an individual spatial effect.
[0230] In some embodiments, in response to detecting a first input, the electronic device presents (1004) a three-dimensional environment with an individual spatial effect at a first immersion level, such as that of FIG. 9B (e.g., presenting the spatial effect within the three-dimensional environment at an initial immersion level) via a display generation component.
[0231] In some embodiments, displaying a spatial effect includes modifying one or more visual characteristics of one or more elements within a three-dimensional environment. As described above, modifying one or more visual characteristics can include changing the illumination temperature, direction, etc. of one or more light sources, which can optionally impart a specific atmosphere effect to the three-dimensional environment and / or objects within the three-dimensional environment, and / or position them at a specific time and / or location, etc. In some embodiments, when the spatial effect is first displayed, it is displayed at an initial immersion level. In some embodiments, the initial immersion level is an intermediate immersion level (e.g., not the maximum immersion level and the minimum immersion level) that can be increased and / or decreased. In some embodiments, the initial immersion level is the minimum immersion level. In some embodiments, the initial immersion level is the maximum immersion level. In some embodiments, a higher immersion level corresponds to an increase in the spatial effect, and a lower immersion level corresponds to a decrease in the spatial effect. For example, a high immersion level for the spatial effect of sunset optionally includes modifying the three-dimensional environment to have a strong orange glow (e.g., the illumination of sunset), and a low immersion level optionally includes modifying the three-dimensional environment to have a gentle orange glow. In some embodiments, displaying a spatial effect includes displaying a simulated environment within the three-dimensional environment. In some embodiments, displaying a simulated environment includes transitioning a portion of the three-dimensional environment to the simulated environment in a manner similar to that described above with respect to method 800. In some embodiments, displaying the simulated environment at a first immersion level includes displaying a subset of the simulated environment (e.g., displaying a first portion of the simulated environment but optionally not displaying a second portion of the simulated environment that is closer to the user).
[0232] In some embodiments, while displaying a three-dimensional environment with individual spatial effects at a first immersion level, the electronic device detects individual inputs (1006) via one or more input devices (e.g., hand movement of one or more hands of the user, movement of one or more eyes of the user and / or change in line of sight, and / or any other input corresponding to a request to change the immersion level of the individual spatial effect) via a hand tracking device, an eye tracking device, or any other input device.
[0233] In some embodiments, the user input includes an operation of a rotational element such as a mechanical dial or a virtual dial (e.g., clockwise rotation to increase, counterclockwise rotation to decrease, or vice versa) to increase or decrease the immersion level. In some embodiments, the user input includes an operation of a displayed control element (e.g., a slider, an immersion indicator, etc.) to increase or decrease the immersion level. In some embodiments, the user input includes a voice input requesting an increase or decrease in the immersion level.
[0234] In some embodiments, in response to detecting an individual input (1008), according to the determination that the individual input is a first input, the electronic device displays, via a display generation component, a three-dimensional environment with individual spatial effects at a second immersion level higher than the first immersion level such as in FIG. 9D (1010) (e.g., increasing the immersion level from an initial immersion level to a second higher immersion level according to the individual input), and according to the determination that the individual input is a second input different from the first input, the electronic device displays, via a display generation component, a three-dimensional environment with individual spatial effects at a third immersion level lower than the first immersion level, such as when the immersion level of the three-dimensional environment 904 in FIG. 9D is 2 or 1 (1012) (e.g., decreasing the immersion level from the initial immersion level to a third lower immersion level according to the individual input).
[0235] For example, if an individual input is a rotation of a rotational element in a direction associated with an increase in the immersion level, the immersion level for the spatial effect is increased to a higher level from its initial position. In some embodiments, increasing the immersion level includes amplifying or increasing the effect of the spatial effect on the three-dimensional environment. For example, if one of the spatial effects for the three-dimensional environment (e.g., one or more spatial effects associated with an individual spatial effect) includes increasing the brightness of the three-dimensional environment, increasing to a higher immersion level from the initial immersion level includes further increasing the brightness of the three-dimensional environment. In some embodiments, increasing the immersion level to a higher level includes increasing the brightness or intensity of one or more visual effects, decreasing the brightness or intensity of one or more visual effects, displaying a visual effect that was not previously displayed, displaying fewer visual effects, changing a visual effect, increasing or decreasing one or more audio effects, playing more audio effects that were not previously played, playing fewer audio effects, changing an audio effect, or any one of the like. In some embodiments, displaying the simulated environment at the second immersion level includes increasing the size of the simulated environment and / or displaying more of the simulated environment (e.g., displaying a first portion of the simulated environment and optionally a second portion of the simulated environment closer to the user).
[0236] For example, if the rotation element rotates in a direction associated with reducing the immersion level of an individual input, the immersion level for the spatial effect decreases to a lower level from its initial position. In some embodiments, reducing the immersion level includes reducing the impact of the spatial effect on the three-dimensional environment. For example, if one of the spatial effects on the three-dimensional environment (e.g., one or more spatial effects associated with an individual spatial effect) includes increasing the brightness of the three-dimensional environment, reducing to a lower immersion level from the initial immersion level includes reducing the brightness of the three-dimensional environment (e.g., optionally, to a level lower than the ambient brightness of the three-dimensional environment without the individual spatial effect, or optionally, to a level equal to or higher than the ambient brightness of the three-dimensional environment without the individual spatial effect). In some embodiments, reducing the immersion level to a higher level includes any one of increasing the brightness or intensity of one or more visual effects, reducing the brightness or intensity of one or more visual effects, displaying a visual effect that was not previously displayed, displaying fewer visual effects, changing the visual effect, increasing or decreasing one or more audio effects, playing more audio effects that were not previously played, playing fewer audio effects, changing the audio effect, etc. In some embodiments, displaying the simulated environment at a third immersion level includes reducing the size of the simulated environment and / or displaying fewer portions of the simulated environment (e.g., displaying fewer portions than the first portion of the simulated environment).
[0237] (For example, by applying an individual spatial effect to a three-dimensional environment at an intermediate immersion level that can be increased or decreased) The above-described method of initiating an individual spatial effect (for example, maintaining the display of at least a portion of the three-dimensional environment and reducing the potentially unpleasant effects of displaying the spatial effect) introduces the spatial effect at a level below the full immersion level, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, (for example, by displaying the spatial effect at an initial immersion level and enabling the user to increase or decrease the magnitude of the spatial effect as desired) makes the user-device interface more efficient, thereby reducing the power consumption further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing the errors during use.
[0238] In some embodiments, the individual spatial effect includes the display of at least a portion of the virtual environment, and displaying a three-dimensional environment without the individual spatial effect includes displaying a representation of a portion of the physical environment around the user's perspective, and displaying a three-dimensional environment with the individual spatial effect includes, as shown in FIG. 9A, replacing at least a portion of the display of the representation of a portion of the physical environment with the display of a portion of the virtual environment (1014) (for example, the spatial effect includes displaying an imitation environment such as that described above with respect to method 800).
[0239] In some embodiments, the imitation environment is displayed in (for example, encompasses) a first portion of the three-dimensional environment. In some embodiments, a second portion of the three-dimensional environment includes a representation of a portion of the physical environment (for example, the real-world environment around the user and / or the device).
[0240] (For example, by initially displaying a simulated environment at an intermediate immersion level that can be increased or decreased) The above-described method of starting the display of a simulated environment (for example, maintaining the display of at least a portion of a three-dimensional environment and reducing the potentially unpleasant effects of displaying a spatial effect) introduces the simulated environment at a level below the full immersion level, (for example, by displaying the simulated environment at an initial immersion level and enabling the user to increase or decrease the amount of the simulated environment as desired) simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing the power consumption further, and improves the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reduces the errors during use.
[0241] In some embodiments, displaying a three-dimensional environment with individual spatial effects includes presenting audio corresponding to a virtual environment, such as that shown by speaker 924 in FIG. 9A (1016) (for example, the individual spatial effects include reproducing audible effects associated with the displayed virtual environment and / or spatial effects (for example, via speakers, via earphones, via headphones, etc.)). In some embodiments, the audible effects include environmental noises or sounds such as running water, wind noise, bird chirping, etc., corresponding to the visual content of the virtual environment.
[0242] In some embodiments, while displaying a three-dimensional environment with individual spatial effects and presenting audio corresponding to the virtual environment, the electronic device detects movement of the electronic device corresponding to changing the orientation of the user's perspective in the three-dimensional environment, such as shown in FIGS. 9E-9H, from a first orientation via one or more input devices (1018) (for example, detecting that the user and / or the device has changed orientation within its physical environment, thereby optionally changing the representation of the physical environment in accordance with the change in the orientation of the user and / or the device).
[0243] For example, when the user rotates the device 45 degrees to the left, the device shifts the display of the three-dimensional environment 45 degrees to the left, exposes the portion of the three-dimensional environment to the left of what was previously displayed, and discontinues the display of a portion of the three-dimensional environment to the right of what was previously displayed. In this way, the environment is adjusted so that it appears as if the user is looking around the three-dimensional environment in the same way the user looks around their physical environment. In some embodiments, when the user and / or the device changes orientation, the user can see the boundary between the physical environment and the virtual environment. In some embodiments, the virtual environment expands in the direction of rotation of the device (e.g., optionally with a delay, following the user's rotation). In some embodiments, the virtual environment does not expand in the direction of rotation of the device.
[0244] In some embodiments, in response to detecting a virtual environment of an electronic device (1020), the electronic device non-emphasizes at least one component of the audio corresponding to the movement, such as in FIG. 9E (1022), according to a determination that the movement of the electronic device meets one or more criteria including a criterion that is met when the orientation of the user's perspective to the three-dimensional environment changes by more than a threshold amount from a first orientation (e.g., when the user and / or the device changes orientation by more than a threshold amount (e.g., more than 30 degrees, 45 degrees, 90 degrees, etc.) or by a predetermined amount (e.g., more than 45 degrees but less than 90 degrees, more than 60 degrees but less than 120 degrees, etc.)), reducing the audible effect provided to the user).
[0245] In some embodiments, reducing the audible effect includes disabling all audible effects associated with spatial effects. In some embodiments, reducing the audible effect includes reducing the volume of the audible effects associated with spatial effects. In some embodiments, the audible effect is reduced and discrete components and / or tracks are removed from the audible effect (e.g., the bird chirping component is removed, but the flowing water component is optionally maintained at the same volume as before).
[0246] (For example, by detecting that the viewing point has changed by a threshold amount and de-emphasizing the audio) the above method of reducing the effect of the audio component enables (for example, automatically reducing the audio effect when the user turns their head, for example, in the direction of the sound source and / or away from the virtual environment, and thus enabling the user to hear sounds from their physical environment without the need to perform additional input to de-emphasize the audio effect) the user to re-engage with their physical environment, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing the power consumption further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing the errors during use.
[0247] In some embodiments, de-emphasizing at least a component of the audio corresponding to the virtual environment includes reducing the degree to which the audio is presented as if it were being generated from different directions within a three-dimensional environment (e.g., reducing the volume of at least one component of the audio associated with the spatial effect), as shown in FIG. 9E (1024). In some embodiments, reducing the degree includes reducing the volume of all components of the audio associated with the spatial effect. In some embodiments, reducing the degree includes reducing the volume of a subset of the components of the audio while maintaining the volume of other components of the audio. In some embodiments, de-emphasizing at least a component of the audio includes reducing the number of audio channels, for example, from five audio channels (e.g., surround sound) to stereo (e.g., two audio channels) and / or from stereo to mono (e.g., one audio channel).
[0248] (For example, by detecting that the viewpoint has changed by a threshold amount and de-emphasizing the audio) the above method of reducing the effect of an audio component enables (for example, automatically reducing the audio effect when the user turns their head, e.g., in the direction of a sound source and / or away from a virtual environment, thus enabling the user to hear sounds from their physical environment without the need to perform additional input to de-emphasize the audio effect) the user to re-engage with their physical environment, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing power consumption further and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing errors during use.
[0249] In some embodiments, de-emphasizing at least a component of the audio corresponding to the virtual environment includes reducing the amount of noise cancellation in which the audio corresponding to the virtual environment is presented (e.g., reducing the noise cancellation effect of the electronic device and / or headphones / speakers / earphones through which the audio is presented) (1026).
[0250] In some embodiments, the electronic device communicates with one or more speakers (or any other noise or audio generating component). In some embodiments, the one or more speakers have one or more noise cancellation features that at least partially reduce (e.g., eliminate) sound from the user's physical environment (e.g., ambient sound). In some embodiments, reducing the noise cancellation effect includes reducing the magnitude of noise cancellation (e.g., canceling 10% less ambient noise, canceling 20% less ambient noise, canceling 50% less ambient noise, etc.). In some embodiments, reducing the noise cancellation effect includes passively enabling ambient noise to reach the user (e.g., disabling noise cancellation). In some embodiments, reducing the noise cancellation effect includes enabling an active transparency listening mode (e.g., to overcome the passive acoustic attenuation effect of headphones and / or earphones, e.g., actively providing ambient sound to the user).
[0251] (E.g., by reducing the noise cancellation function of the device) The above-described method of reducing the immersion effect of the virtual environment provides a quick and efficient way for the user to re-engage with the user's physical environment (e.g., by automatically reducing the audio effect and allowing ambient noise to pass through when the user rotates their head, e.g., in the direction of the sound source and / or away from the virtual environment, enabling the user to hear sound from the user's physical environment without the need for additional input to de-emphasize the audio effect), which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby improving the battery life of the electronic device by further reducing power consumption and enabling the user to use the electronic device more quickly and efficiently, and reducing errors during use.
[0252] In some embodiments, downplaying at least a component of the audio corresponding to the virtual environment includes reducing the number of audio components in which the audio corresponding to the virtual environment is presented (e.g., reducing or eliminating at least a portion of the audio components of the audio associated with the virtual environment) as in FIG. 9E (1028). For example, reducing the volume of one or more tracks of the audio (e.g., to zero) while maintaining the volume of other tracks of the audio.
[0253] The above-described method of reducing the immersive effect of the virtual environment (e.g., by reducing the number of audio components provided) enables the user to re-engage with their physical environment (e.g., by automatically reducing the amount of the audio effect and allowing ambient noise to pass through without the need for the user to perform additional input to downplay the audio effect when the user rotates their head, for example, in the direction of the sound source and / or away from the virtual environment), which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing the power consumption further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing the errors during use.
[0254] In some embodiments, while displaying a three-dimensional environment with individual spatial effects at a first immersion level, the electronic device detects, via one or more input devices, user input corresponding to a request to display a three-dimensional environment with individual spatial effects at a maximum immersion level, such as detecting an input to increase the immersion level to 9 in FIGS. 9E-9F (e.g., receiving user input that interacts with a displayed immersion affordance or interacts with a mechanical dial). In some embodiments, the user input includes a request to increase the immersion to the maximum immersion level. For example, selecting an affordance associated with the highest immersion level and / or performing a rotational input on a mechanical dial associated with requesting the highest immersion level.
[0255] In some embodiments, in response to detecting the user input (1032), according to a determination that a user-defined setting independent of the user input has a first value (e.g., a system setting for setting the maximum immersion level is adjustable by the user), the electronic device displays, via a display generation component, a three-dimensional environment having individual spatial effects at a first individual immersion level (1034) (e.g., if the user-defined setting for the maximum immersion level is the first value (e.g., 180 degrees), increasing the immersion level to the first value based on the user-defined setting), and according to a determination that the user-defined setting has a second value different from the first value, the electronic device displays, via the display generation component, a three-dimensional environment having individual spatial effects at a second individual immersion level greater than the first individual immersion level, such as in FIGS. 9E-9F where the maximum immersion level is bounded by the user-defined setting (1036) (e.g., if the user-defined setting for the maximum immersion level is the second value (e.g., 360 degrees), increasing the immersion level to the second value based on the user-defined setting).
[0256] For example, the maximum immersion level can be 45 degrees of immersion (e.g., a 45-degree three-dimensional environment for the user would be occupied by the virtual object / environment, and the rest would be occupied by the physical environment), 90 degrees of immersion, 180 degrees of immersion, 360 degrees of immersion, etc. In some embodiments, the maximum immersion level can be set to be equal to the field of view (e.g., the maximum angle of immersion is equal to the angle of the field of view of the three-dimensional environment such as 180 degrees or about 180 degrees).
[0257] The above-described method of providing user-defined settings for the maximum immersion level provides a quick and efficient way for the user to be able to limit the maximum immersion level (e.g., by restricting the immersion to a user-defined maximum immersion level without the user having to perform additional user input each time the immersion level is increased to its maximum, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing the power consumption further and improving the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently, and reducing the errors during use).
[0258] In some embodiments, the three-dimensional environment is associated with an application such as FIG. 11D (1038) (e.g., the three-dimensional environment is the user interface of the application). In some embodiments, the application can display a virtual environment. For example, the full-screen mode or presentation mode of the application optionally triggers the display of the virtual environment as described with reference to method 1200.
[0259] In some embodiments, while displaying a three-dimensional environment with individual spatial effects at a first immersion level, the electronic device detects user input (1040) corresponding to a request to display a three-dimensional environment with individual spatial effects at a maximum immersion level via one or more input devices (e.g., receives user input that interacts with a displayed immersion affordance or interacts with a mechanical dial). In some embodiments, the user input includes a request to increase the immersion to the maximum immersion level. For example, selecting an affordance associated with the highest immersion level and / or performing a rotational input on a mechanical dial associated with requesting the highest immersion level.
[0260] In some embodiments, in response to detecting user input (1042), and according to a determination that the application meets one or more first criteria, the electronic device displays, via a display generation component, a three-dimensional environment having individual spatial effects at a first individual immersion level (1044) (e.g., if the application has a maximum immersion level set to the first individual level, displays individual spatial effects at the first individual immersion level), and according to a determination that the application meets one or more second criteria different from the first criteria, the electronic device displays, via the display generation component, a three-dimensional environment having individual spatial effects at a second individual immersion level higher than the first individual immersion level, such as in FIG. 11D (1046) (e.g., if the maximum immersion level set by the application is the second individual level, displays individual spatial effects at the second individual immersion level).
[0261] In some embodiments, the maximum immersion level defined by the application overrides the maximum immersion level of the system settings (e.g., as defined by the user as described above). For example, if the maximum immersion level for the application is greater than the maximum immersion level of the system, the maximum immersion level for the application controls. In some embodiments, the maximum immersion level is limited by the immersion level of the system even if the maximum immersion level of the application is greater. In some embodiments, if the maximum immersion level defined for the application is less than the immersion level of the system, in response to a request to display a spatial effect at the maximum immersion level, the spatial effect is displayed at the maximum level defined by the application. Thus, in some embodiments, the immersion level is the lesser of the maximum level set by the application or the maximum level set by the system level settings.
[0262] The above-described method of determining the maximum immersion level based on the level set by the application (e.g., by defining the maximum immersion level based on the level set by the active application) simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing the power consumption further and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing the errors during use.
[0263] In some embodiments, an individual spatial effect includes an ambient effect (e.g., a visual modification of at least a portion of a three-dimensional environment (e.g., not associated with an object within the three-dimensional environment)), and displaying a three-dimensional environment without an individual spatial effect includes displaying a representation of a portion of the physical environment around the user's viewpoint without an ambient effect, and displaying a three-dimensional environment with an individual spatial effect includes displaying a representation of a portion of the physical environment around the user's viewpoint with an ambient effect, as shown in FIG. 9E for example (1048) (e.g., the ambient effect is displayed within a portion of the three-dimensional environment that includes the spatial effect, but not within a portion of the three-dimensional environment that does not include the spatial effect).
[0264] For example, display an ambient lighting effect (e.g., sunrise, sunset, etc.), a fog effect, a mist effect, a smoke / particle effect, etc. In some embodiments, the ambient effect is such that the air or empty space of the three-dimensional environment appears to be filled with a physical effect.
[0265] For example, if the level is immersion, 30% of the "room" shows a spatial effect and 70% of the "room" does not show a spatial effect (e.g., optionally a representation of a room not modified by any spatial effect), 30% of the room with the spatial effect is displayed with an atmosphere like fog, and 70% of the room without the spatial effect does not include an atmosphere like fog. In some embodiments, showing a spatial effect includes modifying a portion of the physical environment to have a spatial effect (e.g., modifying the lighting of an object within the physical environment, modifying the visual characteristics of an object within the physical environment, etc.).
[0266] (For example, by displaying an ambient effect in a portion of the environment where a spatial effect is displayed but not in a portion of the environment that does not include the spatial effect) The above-described method of displaying a spatial effect provides a quick and efficient way to display a spatial effect in a subset of the environment, which simplifies the interaction between the user and the electronic device, improves the operability of the electronic device, makes the user-device interface more efficient, thereby reducing the power consumption further and enabling the user to use the electronic device more quickly and efficiently, improving the battery life of the electronic device and reducing the errors during use.
[0267] In some embodiments, displaying a three-dimensional environment with individual spatial effects includes presenting audio corresponding to the ambient effect (1050) as shown in FIG. 9B (for example, an individual spatial effect includes playing an audible effect associated with the displayed virtual environment and / or the spatial effect). In some embodiments, the audible effect includes noise or sounds associated with the ambient effect, such as the sound of running water, wind noise, bird chirping, the crackling sound of a burning tree, etc.
[0268] In some embodiments, while displaying a three-dimensional environment with individual spatial effects and presenting audio corresponding to the ambient effect, the electronic device detects movement of the electronic device corresponding to changing the orientation of the user's perspective in the three-dimensional environment, such as in FIGS. 9E-9H, from a first orientation via one or more input devices (1052) (for example, detecting that the user and / or the device has changed its orientation within its physical environment, thereby optionally changing the representation of the physical environment according to the ch...
Claims
Claim 1 In an electronic device that communicates with a display generation component and one or more input devices, while an object corresponding to the perspective of a user of the electronic device is at a first location within an area in the physical environment of the electronic device, displaying, via the display generation component, a three-dimensional environment from the perspective of the user, the three-dimensional environment including a simulated environment and a boundary between a first portion of the three-dimensional environment that includes the simulated environment and a second portion of the three-dimensional environment, the boundary being at a first distance from the perspective of the user within the three-dimensional environment; while displaying the three-dimensional environment including the simulated environment and the boundary at the first distance from the perspective of the user, detecting, via the one or more input devices, a movement of the object corresponding to the perspective of the user from the first location to a second location within the physical environment; in response to detecting the movement of the object corresponding to the perspective of the user from the first location to the second location; updating the three-dimensional environment to change an amount of the three-dimensional environment occupied by the simulated environment such that the boundary of the simulated environment is at a second distance greater than the first distance from the perspective of the user, according to a determination that the movement of the object corresponding to the perspective of the user from the first location to the second location satisfies one or more criteria, wherein updating the three-dimensional environment includes replacing a display of a portion of the simulated environment with a display of a representation of a portion of the physical environment of the electronic device; A method comprising the above. Claim 2 The boundary between the first portion and the second portion of the three-dimensional environment includes a region between the simulated environment and the second portion of the three-dimensional environment, wherein at a particular location within the region, at least an individual portion of the simulated environment and at least an individual portion of the second portion of the three-dimensional environment are simultaneously displayed; The method according to claim 1. Claim 3 Reducing the immersion level at which the simulated environment is displayed in response to detecting the movement of the object corresponding to the user's perspective from the first location to the second location and according to the determination that the movement of the object corresponding to the user's perspective from the first location to the second location in the physical environment meets the one or more criteria. The method according to claim 1, further comprising.
4. Updating the three-dimensional environment to change the amount of the three-dimensional environment occupied by the simulated environment such that the boundary of the simulated environment is at the second distance from the user's perspective, which is independent of whether there are obstacles in the physical environment of the electronic device. The method according to claim 1.
5. After detecting the movement of the object corresponding to the user's perspective from the first location to the second location, while displaying the three-dimensional environment including the simulated environment and the boundary at the second distance from the user's perspective, detecting a second movement of the object corresponding to the user's perspective away from the second location in the physical environment via the one or more input devices; In response to detecting the second movement of the object corresponding to the user's perspective away from the second location in the physical environment, Updating the three-dimensional environment to change the amount of the three-dimensional environment occupied by the simulated environment, including replacing the representation of the portion of the physical environment with a portion of the simulated environment according to the determination that the second movement of the object corresponding to the user's perspective is away from the boundary; The method according to claim 1, further comprising.
6. Updating the three-dimensional environment to change the amount of the three-dimensional environment occupied by the simulated environment is According to the determination that the second location is at a first individual distance from the boundary, replacing the display of the first amount of the simulated environment with the display of the first amount of the physical environment of the electronic device. In accordance with the determination that the second location is a second individual distance different from the first individual distance from the boundary, replacing a display of a second amount different from the first amount in the simulated environment with a display of the second amount in the physical environment of the electronic device. The method according to claim 1, comprising:
7. Updating the three-dimensional environment to change the amount of the three-dimensional environment occupied by the simulated environment, In accordance with the determination that the speed of the movement of the object corresponding to the user's perspective is a first speed, replacing a display of a first amount in the simulated environment with a display of the first amount in the physical environment of the electronic device; In accordance with the determination that the speed of the movement of the object corresponding to the user's perspective is a second speed different from the first speed, replacing a display of a second amount different from the first amount in the simulated environment with a display of the second amount in the physical environment of the electronic device. The method according to claim 1, comprising:
8. Updating the three-dimensional environment to change the amount of the three-dimensional environment occupied by the simulated environment, In accordance with the determination that the acceleration of the movement of the object corresponding to the user's perspective is a first acceleration, replacing a display of a first amount in the simulated environment with a display of the first amount in the physical environment of the electronic device; In accordance with the determination that the acceleration of the movement of the object corresponding to the user's perspective is a second acceleration different from the first acceleration, replacing a display of a second amount different from the first amount in the simulated environment with a display of the second amount in the physical environment of the electronic device. The method according to claim 1, comprising:
9. After detecting the movement of the object corresponding to the user's perspective from the first location to the second location, while displaying the three-dimensional environment including the simulated environment and the boundary at the second distance from the user's perspective, detecting, via the one or more input devices, a second movement of the object corresponding to the user's perspective away from the second location in the physical environment; In response to detecting the second movement of the object corresponding to the user's viewpoint away from the second location within the physical environment, updating the three-dimensional environment to stop displaying the imitation environment within the three-dimensional environment according to a determination that the second movement of the object corresponding to the user's viewpoint is towards the boundary; The method according to claim 1, further comprising.
10. While displaying the three-dimensional environment including the imitation environment and the boundary at the first distance from the user's viewpoint, via the one or more input devices, detecting a change in the orientation of the object corresponding to the user's viewpoint without detecting a movement of the object corresponding to the user's viewpoint away from the first location within the physical environment; In response to detecting the change in the orientation of the object corresponding to the user's viewpoint, while maintaining the boundary at the first distance from the user's viewpoint, updating the display of the three-dimensional environment according to the change in the orientation of the object corresponding to the user's viewpoint without replacing the display of the portion of the imitation environment with the display of the portion of the physical environment of the electronic device; The method according to claim 1, further comprising.
11. The method according to claim 1, wherein the one or more criteria are satisfied when the second location is within a threshold distance of the location corresponding to the boundary and not satisfied when the second location is farther than the threshold distance from the location corresponding to the boundary.
12. Updating the three-dimensional environment to change the amount of the three-dimensional environment occupied by the imitation environment includes changing which portion of the three-dimensional environment is occupied by the imitation environment based on the distance of one or more physical characteristics within the physical environment of the electronic device from the object corresponding to the user's viewpoint. The method according to claim 1.
13. Displaying a representation of a third portion of the physical environment with an ambient effect, while displaying the three-dimensional environment, detecting a second movement of an object corresponding to the user's view point from the first location to the second location within the physical environment via the one or more input devices, Updating the display of the three-dimensional environment in accordance with the second movement of the object corresponding to the user's view point, without changing the ambient effect in which the representation of the third portion of the physical environment is displayed, in response to detecting the second movement of the object corresponding to the user's view point from the first location to the second location, the method according to claim 1, further comprising.
14. Before detecting the movement of the object corresponding to the user's view point from the first location to the second location, the three-dimensional environment includes a user interface of an application located within the imitation environment, and the method In response to detecting the movement of the object corresponding to the user's view point from the first location to the second location, according to the determination that the movement of the object corresponding to the user's view point from the first location to the second location within the physical environment meets one or more second criteria Updating the three-dimensional environment so as not to include the imitation environment any more The method according to claim 1, further comprising stopping the display of the user interface of the application in the three-dimensional environment.
15. Before detecting the movement of the object corresponding to the user's view point from the first location to the second location, the three-dimensional environment includes a user interface of an application located outside the imitation environment within the three-dimensional environment, and the method In response to detecting the movement of the object corresponding to the user's view point from the first location to the second location, according to the determination that the movement of the object corresponding to the user's view point from the first location to the second location within the physical environment meets one or more second criteria Updating the three-dimensional environment so as to no longer include the emulation environment; The method according to claim 1, further comprising maintaining a display of the user interface of the application in the three-dimensional environment.
16. Before detecting the movement of the object corresponding to the user's viewpoint from the first location to the second location, the three-dimensional environment includes a user interface of an application located within the emulation environment, and the method In response to detecting the movement of the object corresponding to the user's viewpoint from the first location to the second location, according to a determination that the movement of the object corresponding to the user's viewpoint from the first location to the second location in the physical environment satisfies one or more second criteria, Updating the first portion of the three-dimensional environment so as to no longer include the emulation environment; The method according to claim 1, further comprising displaying the user interface of the application in the second portion of the three-dimensional environment.
17. While the user interface of the application is located in the first portion of the three-dimensional environment, the user interface of the application has a first size in the three-dimensional environment, While the user interface of the application is located in the second portion of the three-dimensional environment, the user interface of the application has a second size smaller than the first size in the first portion of the three-dimensional environment, the method according to claim 16.
18. One or more processors; A memory; One or more programs; An electronic device comprising: The one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 1 to 17.
19. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs comprise instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Head mounted display device and image display system
JP2013178639A
Anti-stumbling when immersed in a virtual reality environment
JP2017531221A
Display unit and method for controlling display unit
JP2018101019A
Monitor motion sickness and provide additional sounds to reduce motion sickness
JP2018514005A
Detecting the end of a session in an augmented reality and / or virtual reality environment
JP2020503595A