Method for improving user's environmental perception
By improving the user interface and interaction methods, the complex and power consumption of virtual/augmented reality environment interaction in the prior art is solved, and a more efficient and intuitive user interaction experience is achieved, and battery power is saved.
Patent Information
- Application Number
- CN202380067876.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-04
- Filing Date
- 2023-09-23
- Publication Date
- 2025-05-06
AI Technical Summary
When interacting with virtual/augmented reality environments, existing computer systems lack feedback, complex operations, cumbersome and prone to errors, resulting in large cognitive burden on users, low interaction efficiency, and power consumption.
By providing improved user interfaces and interaction methods, the number and complexity of user input is reduced and the efficiency of the human-computer interface is improved. Specific measures include displaying interactive areas of virtual content, generating warnings for physical objects, reducing visual salience of virtual content, and attention-based interaction methods.
It improves the interaction efficiency and intuitiveness between users and computer systems, reduces user errors and cognitive burden, saves battery power, and extends the use time of the device.
Smart Images

Figure CN119948437A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 376,961, filed on September 23, 2022, and U.S. Provisional Application No. 63 / 506,095, filed on June 4, 2023, the contents of both provisional applications are incorporated herein by reference in their entirety for all purposes. Technical Field
[0003] The present disclosure generally relates to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via displays. Background Art
[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touch screen displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the invention
[0005] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems that make virtual object manipulation complex, cumbersome, and error-prone can impose a huge cognitive burden on users and detract from the experience of virtual / augmented reality environments. In addition, these methods take longer than necessary, wasting energy on the computer system. This latter consideration is particularly important in battery-powered devices.
[0006] Therefore, there is a need for computer systems with improved methods and interfaces to provide computer-generated experiences to users, so that the user's interaction with the computer system is more efficient and more intuitive to the user. Such methods and interfaces optionally supplement or replace conventional methods for providing extended reality experiences to users. Such methods and interfaces reduce the amount, degree and / or nature of inputs from the user by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby forming a more effective human-computer interface.
[0007] The above-mentioned defects and other problems associated with the user interface of the computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or a head-mounted device). In some embodiments, the computer system has a touch pad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye tracking components. In some embodiments, the computer system has one or more hand tracking components. In some embodiments, in addition to the display generation component, the computer system also has one or more output devices, which include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or instruction set stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other mobile sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed by interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, testing support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] There is a need for electronic devices with improved methods and interfaces for interacting with content in a three-dimensional environment. Such methods and interfaces can supplement or replace conventional methods for interacting with content in a three-dimensional environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from a user and produce a more efficient human-computer interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.
[0009] In some embodiments, the computer system displays virtual content showing areas of possible interaction and displays immersive virtual content. In some embodiments, the computer system stops display of immersive virtual content and displays areas of possible interaction. In some embodiments, the computer system generates warnings for physical objects obscured by virtual content based on attention. In some embodiments, the computer system reduces the visual salience of virtual content for people in the physical environment based on attention.
[0010] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not comprehensive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art based on the drawings, the specification, and the claims. In addition, it should be noted that the language used in this specification is selected in principle for readability and instructional purposes, and may not be selected to describe or define the subject matter of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the drawings.
[0012] Figure 1A is a block diagram illustrating an operating environment of a computer system for providing an XR experience according to some embodiments.
[0013] Figure 1B to Figure 1P is used in Figure 1A An example of a computer system that provides an XR experience in an operating environment.
[0014] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.
[0015] Figure 3 is a block diagram illustrating display generation components of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.
[0016] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture input from a user according to some embodiments.
[0017] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture gaze input of a user according to some embodiments.
[0018] Figure 6 is a flow chart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.
[0019] FIG. 7A to FIG. 7D An example of a computer system that displays virtual content showing areas of possible interaction and displays immersive virtual content according to some embodiments is shown.
[0020] FIG. 8A to FIG. 8F is a flowchart illustrating an exemplary method of displaying virtual content showing areas of possible interaction and displaying immersive virtual content according to some embodiments.
[0021] 9A to 9E An example of a computer system that reduces the visual prominence of immersive virtual content and displays areas of possible interaction according to some embodiments is shown.
[0022] FIG. 10A to FIG. 10G is a flowchart illustrating a method of reducing the visual prominence of immersive virtual content and displaying areas of possible interaction according to some embodiments.
[0023] FIG. 11A to FIG. 11E An example of a computer system that generates alerts associated with physical objects in a user's environment according to some embodiments is shown.
[0024] FIG. 12A to FIG. 12D is a flowchart illustrating a method of generating alerts associated with physical objects in a user's environment according to some embodiments.
[0025] FIG. 13A to FIG. 13H An example of a computer system that alters the visual saliency of a person in a three-dimensional environment based on one or more attention-related factors according to some embodiments is shown.
[0026] FIG. 14A to FIG. 14H is a flowchart illustrating a method for changing the visual saliency of a person in a three-dimensional environment based on one or more attention-related factors according to some embodiments. DETAILED DESCRIPTION
[0027] According to some embodiments, the present disclosure relates to a user interface for providing a computer-generated (CGR) experience to a user.
[0028] The systems, methods, and GUIs described herein provide improved ways for electronic devices to facilitate interaction with and manipulation of objects in a three-dimensional environment.
[0029] In some embodiments, a computer system detects an input corresponding to a request to display virtual content at an immersion level greater than a threshold immersion level. In some embodiments, the computer system displays a visual indication corresponding to an area where interaction with the virtual content is possible. In some embodiments, the input includes moving a user of the computer system into an area of the user's physical environment corresponding to the visual indication. In some embodiments, the computer system maintains display of a portion of a representation of the user's environment while displaying the virtual content at an immersion level greater than the threshold immersion level.
[0030] In some embodiments, the computer system detects an input corresponding to a request to reduce the visual prominence of the virtual content. In some embodiments, the input includes moving a user of the computer system out of an area of the user's physical environment where the computer system expects that the user may interact with the virtual content. In some embodiments, reducing the visual prominence includes stopping display of the virtual content.
[0031] In some embodiments, the computer system displays virtual content that obscures physical objects in the user's physical environment. In some embodiments, based on determining that the physical object may conflict with the user's range of motion, the computer system generates a warning indicating the presence of the physical object. In some embodiments, based on the user's attention to the warning, the computer system reduces, maintains, or increases the prominence of the warning.
[0032] In some embodiments, the computer system displays virtual content that obscures a person in the physical environment of the computer system. In some embodiments, the computer system breaks through the virtual content to allow visibility of the person through the virtual content. In some embodiments, the computer system changes the visibility of the person through the virtual content based on the user and / or the person's attention.
[0033] Figures 1A to 6 A description of an example computer system for providing an XR experience to a user (such as described below with respect to methods 800 , 1000 , 1200 , and / or 1400 ) is provided. FIG. 7A to FIG. 7D An example of a computer system that displays virtual content showing areas of possible interaction and displays immersive virtual content according to some embodiments is shown. FIG. 8A to FIG. 8F is a flowchart illustrating an exemplary method of displaying virtual content showing areas of possible interaction and displaying immersive virtual content according to some embodiments. FIG. 7A to FIG. 7D The user interface in FIG. 8A to FIG. 8F process. 9A to 9E An example of a computer system that reduces the visual prominence of immersive virtual content and displays areas of possible interaction according to some embodiments is shown. FIG. 10A to FIG. 10Gis a flowchart illustrating a method of reducing the visual prominence of immersive virtual content and displaying areas of possible interaction according to some embodiments. 9A to 9E The user interface in FIG. 10A to FIG. 10G process. FIG. 11A to FIG. 11E Example techniques for generating alerts associated with physical objects in a user's environment are shown according to some embodiments. FIG. 12A to FIG. 12D is a flow diagram of a method of generating an alert associated with a physical object in a user's environment according to various embodiments. FIG. 11A to FIG. 11E The user interface in FIG. 12A to FIG. 12D process. FIG. 13A to FIG. 13H Example techniques are shown for changing the visual saliency of a person in a three-dimensional environment based on one or more attention-related factors according to some embodiments. FIG. 14A to FIG. 14H is a flow chart of a method for changing the visual saliency of a person in a three-dimensional environment based on one or more attention-related factors according to various embodiments. FIG. 13A to FIG. 13H The user interface in FIG. 14A to FIG. 14H process.
[0034] The process described below enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating the device / interacting with the device) through various technologies, including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation without further user input when a set of conditions have been met, improving privacy and / or security, providing a more diverse, detailed and / or realistic user experience while saving storage space, and / or additional technologies. These technologies also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. Saving battery power, and therefore saving weight, improves the ergonomics of the device. These technologies also enable real-time communication, allow the use of fewer and / or less accurate sensors, thereby producing a more compact, lighter and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage and thereby reduce the heat emitted by the device, which is particularly important for wearable devices where if the device generates too much heat well within the operating parameters of the device components, it may become uncomfortable for the user to wear the device.
[0035] In addition, in the method described herein where one or more steps depend on one or more conditions being met, it should be understood that the method can be repeated in multiple repetitions so that in the process of repetition, all conditions for determining the steps in the method have been met in different repetitions of the method. For example, if the method requires performing the first step (if the condition is met), and performing the second step (if the condition is not met), then the skilled person will know that the steps stated are repeated until both the condition is met and the condition is not met (in no particular order). Therefore, the method described as having one or more steps depending on one or more conditions being met can be rewritten as a method of repeating until each condition described in the method is met. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing a contingent operation based on the satisfaction of the corresponding one or more conditions, and is therefore able to determine whether a possible situation has been met without explicitly repeating the steps of the method until all conditions for determining the steps in the method have been met. It will also be understood by those of ordinary skill in the art that, similar to the method with the contingent steps, the system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all the contingent steps have been performed.
[0036] In some embodiments, such as Figure 1A As shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a household appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head mounted device or a handheld device).
[0037] In describing an XR experience, various terms are used to distinguishably refer to several related but distinct environments that a user can sense and / or with which the user can interact (e.g., using inputs detected by the computer system 101 generating the XR experience, which inputs cause the computer system generating the XR experience to generate audio, visual, and / or tactile feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:
[0038] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of electronic systems. A physical environment, such as a physical park, includes physical objects, such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.
[0039] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic system. In XR, a subset of a person's physical movements, or a representation thereof, is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. For example, an XR system may detect a person's head turn, and in response, adjust the graphical content and sound field presented to the person in a manner similar to the way such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), adjustments to the characteristics of virtual objects in the XR environment may be made in response to a representation of physical movement (e.g., a voice command). People can sense and / or interact with XR objects using any of their senses, including vision, hearing, touch, taste, and smell. For example, people can sense and / or interact with audio objects, which create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with an audio object.
[0040] Examples of XR include virtual reality and mixed reality.
[0041] Virtual Reality: A virtual reality (VR) environment refers to a simulated environment designed to be based entirely on computer-generated sensory input to one or more senses. A VR environment includes a plurality of virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing avatars of people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of the person's presence within the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.
[0042] Mixed Reality: In contrast to VR environments, which are designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment refers to a simulated environment that is designed to include sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtuality continuum, a mixed reality environment is anything between, but not including, a fully physical environment at one end and a virtual reality environment at the other end. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. In addition, some electronic systems used to render MR environments can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical items from the physical environment, or representations thereof). For example, the system can cause motion so that virtual trees appear stationary relative to the physical ground.
[0043] Examples of mixed reality include augmented reality and augmented virtuality.
[0044] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on a transparent or translucent display so that a person uses the system to perceive virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the image or video with the virtual object and presents the composition on an opaque display. A person uses the system to indirectly view the physical environment via an image or video of the physical environment and perceives virtual objects superimposed on the physical environment. As used herein, a video of a physical environment displayed on an opaque display is referred to as a "transparent video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on an opaque display. Further alternatively, the system may have a projection system that projects virtual objects into a physical environment, such as as a hologram or on a physical surface, so that a person using the system perceives virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a pass-through video, the system may transform one or more sensor images to apply a selected perspective (e.g., a viewpoint) that is different from the perspective captured by the imaging sensor. For another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion thereof so that the modified portion may be a representative but not real version of the original captured image. For another example, the representation of the physical environment may be transformed by graphically eliminating a portion thereof or blurring a portion thereof.
[0045] Augmented Virtual: An augmented virtual (AV) environment refers to a simulated environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input may be a representation of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but the faces of people are realistically reproduced from images taken of physical people. For another example, a virtual object may take the shape or color of a physical object imaged by one or more imaging sensors. For another example, a virtual object may take a shadow that conforms to the positioning of the sun in the physical environment.
[0046] In augmented reality, mixed reality or virtual reality environment, the view of the three-dimensional environment is visible to the user. The view of the three-dimensional environment is usually visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), and the virtual viewport has a viewport boundary, which defines the scope of the three-dimensional environment visible to the user via one or more display generation components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual range in one or more dimensions (e.g., based on the user's visual range, the size of one or more display generation components, optical properties or other physical properties, and / or the position and / or orientation of one or more display generation components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual range in one or more dimensions (e.g., based on the user's visual range, the size of one or more display generation components, optical properties or other physical properties, and / or the position and / or orientation of one or more display generation components relative to the user's eyes). The viewport and viewport boundary usually move with the movement of one or more display generation components (e.g., for head-mounted devices, they move with the user's head, or for handheld devices such as tablet computers or smart phones, they move with the user's hands). The user's viewpoint determines what is visible in the viewport, and the viewpoint typically specifies a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a view of the three-dimensional environment that is perceptually accurate and provides an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint moves as the handheld or fixed device moves and / or as the user's positioning relative to the handheld or fixed device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For devices including display generation components with virtual pass-through, portions of the physical environment that are visible (e.g., displayed and / or projected) via one or more display generation components are based on the field of view of one or more cameras in communication with the display generation components, which one or more cameras typically move with movement of the display generation components (e.g., with movement of the user's head for a head-mounted device, or with movement of the user's hands for a handheld device such as a tablet computer or smart phone) because the user's viewpoint moves with movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on movement of the user's viewpoint)).For display generating components with optical transmittance, portions of the physical environment that are visible via one or more display generating components (e.g., optically visible through one or more partially or fully transparent portions of the display generating components) are based on the user's field of view through the partially or fully transparent portions of the display generating components (e.g., moves as the user's head moves for a head-mounted device, or moves as the user's hands move for a handheld device such as a tablet computer or smart phone) because the user's viewpoint moves as the user moves through the field of view of the partially or fully transparent portions of the display generating components (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0047] In some embodiments, the representation of the physical environment (e.g., displayed via virtual transmission or optical transmission) may be partially or completely obscured by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment that is not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or obscuring more of the physical environment, and reducing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or obscured. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are more visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes an associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by a computer system obscures background content (e.g., content other than the virtual environment and / or virtual content) surrounding / behind the virtual environment, optionally including the number of items of background content displayed and / or displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generation component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generation component that is occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application), virtual objects that are not associated with or included in the virtual environment and / or virtual content (e.g., files generated by a computer system or representations of other users, etc.), and / or real objects (e.g., transparent objects representing real objects in the physical environment surrounding the user, which are visible so that they are displayed via the display generation component and / or visible via transparent or translucent components of the display generation component because the computer system does not block / impede their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with the background content, which is optionally displayed at full brightness, color and / or translucency.In some embodiments, at a higher immersion level (e.g., a second immersion level higher than the first immersion level), background, virtual and / or real objects are displayed in an obstructed manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full screen or full immersion mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary between background objects. For example, at a specific immersion level, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects stop displaying. In some embodiments, zero immersion or zero immersion level corresponds to a virtual environment that stops displaying, and instead displays a representation of the physical environment (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), while the representation of the physical environment is not obscured by the virtual environment. Adjusting the immersion level using physical input elements provides a fast and efficient method of adjusting the degree of immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.
[0048] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or location in a user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments where the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is an augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or location of the viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."
[0049] Environment-locked visual objects: A virtual object is environment-locked (alternatively, "world-locked") when a computer system displays a virtual object at a position and / or location in a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint moves, the location and / or objects in the environment change relative to the user's viewpoint, which causes the environment-locked virtual object to be displayed at a different location and / or location in the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user is displayed at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center in the user's viewpoint (e.g., the tree location in the user's viewpoint shifts), the environment-locked virtual object locked to the tree is displayed to the left of center in the user's viewpoint. In other words, the position and / or orientation of the virtual object that is locked to the environment in the user's viewpoint depends on the position and / or orientation of the object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference system (e.g., a coordinate system anchored to a fixed position and / or object in the physical environment) to determine the position of the virtual object that is locked to the user's viewpoint. The virtual object that is locked to the environment can be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object), or can be locked to a movable part of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot of the user) so that the virtual object moves as the viewpoint or that part of the environment moves to maintain a fixed relationship between the virtual object and that part of the environment.
[0050] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits an inertial following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of the reference point followed by the virtual object. In some embodiments, when exhibiting inertial following behavior, when detecting the movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5cm and 300cm from the viewpoint) that the virtual object is following, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., the portion of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits inertial following behavior, the device ignores small amounts of movement of the reference point (e.g., ignoring movements of the reference point below a threshold movement amount, such as moving 0 degrees to 5 degrees or moving 0cm to 50cm). For example, when a reference point (e.g., a portion or viewpoint of an environment to which a virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., the portion or viewpoint of the environment to which the virtual object is locked) moves a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a “lazy follow” threshold) because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some embodiments, maintaining a substantially fixed position of the virtual object relative to the reference point includes displaying the virtual object within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the position of the reference point).
[0051] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Instead of an opaque display, a head-mounted system can have a transparent or translucent display. A transparent or translucent display can have a medium through which light representing the image is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to become selectively opaque. The projection-based system can employ retinal projection technology that projects graphic images onto a person's retina. The projection system can also be configured to project virtual objects into a physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The following description with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is located locally or remotely relative to scene 105 (e.g., physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., cloud server, central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., HMD, display, projector, touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is included within a housing (e.g., a physical housing) of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices of the input devices 125, one or more output devices of the output devices 155, one or more sensors of the sensors 190, and / or one or more peripheral devices of the peripheral devices 195, or shares the same physical housing or support structure with one or more of the above devices.
[0052] In some embodiments, the display generation component 120 is configured to provide an XR experience (e.g., at least the visual component of the XR experience) to the user. In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Figure 3 Display generation component 120 is described in further detail. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0053] According to some embodiments, display generation component 120 provides an XR experience to the user when the user is virtually and / or physically present within scene 105.
[0054] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). In this way, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or tablet device) configured to present XR content, and the user holds a device with a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a shell worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR room, shell, or room configured to present XR content, wherein the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in the space in front of a handheld device or a tripod-mounted device may similarly be implemented with an HMD, where the interactions occur in the space in front of the HMD and responses to the XR content are displayed via the HMD. Similarly, a user interface showing interactions with XR content triggered based on movement of a handheld device or a tripod-mounted device relative to a physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)) may similarly be implemented with an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)).
[0055] Despite Figure 1A Relevant features of operating environment 100 are shown, but those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example embodiments disclosed herein.
[0056] Figure 1A to Figure 1PVarious examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interface described herein are shown. In some embodiments, the computer system includes one or more display generation components (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) for displaying to a user of the computer system a representation of a virtual element and / or a physical environment, optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, which are optionally removably attached to one or more of the optical modules to enable the user interface to be more easily viewed by users who would otherwise use glasses or contact lenses to correct their vision. Although many of the user interfaces shown herein show a single view of the user interface, the user interface in the HMD is optionally displayed using two optical modules (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b), one optical module for the user's right eye and a different optical module for the user's left eye, and slightly different images are presented to the two different eyes to generate the illusion of stereoscopic depth, the single view of the user interface is typically a right eye view or a left eye view, and the depth effect is explained in the text or using other schematics or views. In some embodiments, the computer system includes one or more external displays (e.g., display components 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not worn) and / or to other people near the computer system, the status information being optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., sensor components 1-356 and / or Fig. 1I ), which information can be used (optionally in conjunction with one or more illuminators, such as Fig. 1IIn some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand position and / or movement (e.g., sensor assembly 1-356 and / or sensor assembly 1-357 and / or sensor assembly 1-358 and / or sensor assembly 1-359). Fig. 1I One or more sensors in the apparatus) which may be used (optionally in combination with one or more illuminators, such as Fig. 1I In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Fig. 1I eye tracking and gaze tracking sensors in the Fig.1OLights 11.3.2-110 in the device) determine attention or gaze location and / or gaze movement, which can be optionally used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for a real-time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. Gaze and / or attention information is optionally combined with hand tracking information to determine interaction between a user and one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), knobs (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328), digital crown (e.g., pressable and twistable or rotatable first button 1-128, button 11.1.1-114, and / or dial or button 1-328), touchpad, touch screen, keyboard, mouse, and / or other input devices. One or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328) are optionally used to perform system operations, such as re-centering content in a three-dimensional environment visible to a user of the device, displaying a primary user interface for launching an application, starting a real-time communication session, or initiating display of a virtual three-dimensional background. A knob or digital crown (e.g., a pressable and twistable or rotatable first button 1-128, button 11.1.1-114, and / or dial or button 1-328) is optionally rotatable to adjust parameters of visual content, such as an immersion level of a virtual three-dimensional environment (e.g., the extent to which virtual content occupies a user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and virtual content displayed via optical modules (e.g., first display component 1-120a and second display component 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).
[0057] Figure 1BFront, top, and perspective views of an example of a head-mountable display (HMD) device 1-100 configured to be worn by a user and to provide a virtual and altered / mixed reality (VR / AR) experience are shown. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strap assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured to the electronic strap assembly 1-104 at either end. The electronic strap assembly 1-104 and the strap 1-106 may be part of a retaining assembly configured to wrap around a user's head to hold the display unit 1-102 against the user's face.
[0058] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of a user's head and a second strap 1-117 configured to extend over the top of the user's head. As shown, the second strap may extend between the first electronic strip 1-105a and the second electronic strip 1-105b of the electronic strip assembly 1-104. The strap assembly 1-104 and the strap assembly 1-106 may be part of a securing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.
[0059] In at least one example, the securing mechanism includes a first electronic strip 1-105a including a first proximal end 1-134 coupled to the display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The securing mechanism may also include a second electronic strip 1-105b including a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The securing mechanism may also include a first band 1-116 and a second band 1-117, the first band including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second band extending between the first electronic strip 1-105a and the second electronic strip 1-105b. The strips 1-105a-b and the strip 1-116 may be coupled via a connection mechanism or assembly 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strip 1-105b between the second proximal end 1-138 and the second distal end 1-140.
[0060] In at least one example, the first and second electronic strips 1-105a-b include plastic, metal, or other structural materials formed into substantially rigid strip 1-105a-b shapes. In at least one example, the first band 1-116 and the second band 1-117 are formed of a resilient flexible material including a woven textile, rubber, etc. The first band 1-116 and the second band 1-117 can be flexible to conform to the shape of the user's head when the HMD 1-100 is worn.
[0061] In at least one example, one or more of the first and second electronic strips 1-105a-b can define an inner strip volume and include one or more electronic components disposed in the inner strip volume. Figure 1B As shown, the first electronic strip 1-105a may include an electronic component 1-112. In one example, the electronic component 1-112 may include a speaker. In one example, the electronic component 1-112 may include a computing component, such as a processor.
[0062] In at least one example, the housing 1-150 defines a first front opening 1-152. Figure 1B 1-152 in dashed lines because the display assembly 1-108 is configured to shield the first opening 1-152 from the field of view when the HMD 1-100 is assembled. The housing 1-150 may also define a rear-mounted second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover and a display screen (shown in other figures) disposed in or across the front opening to shield the front opening 1-152. In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 generally, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 may be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, where the display unit 1-102 is pressed.
[0063] In at least one example, the housing 1-150 may define a first hole 1-126 between the first opening 1-152 and the second opening 1-154 and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 can be pressed through the corresponding holes 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twistable dial and a pressable button. In at least one example, the first button 1-128 is a pressable and twistable dial button, and the second button 1-132 is a pressable button.
[0064] Figure 1C A rear perspective view of the HMD 1-100 is shown. The HMD 1-100 may include a light seal 1-110 extending rearwardly from the housing 1-150 of the display assembly 1-108 around the periphery of the housing 1-150, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to the face of the user, surrounding the eyes of the user, to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b, which are disposed at or in a rearward-facing second opening 1-154 defined by the housing 1-150 and / or disposed in the interior volume of the housing 1-150 and are configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a corresponding display screen 1-122a, 1-122b, which is configured to project light in a rearward direction through the second opening 1-154 toward the eyes of the user.
[0065] In at least one example, reference Figure 1B and Figure 1C In both cases, the display assembly 1-108 may be a front-facing forward display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b may be configured to project light in a second rearward direction opposite the first direction. As described above, the light seal 1-110 may be configured to block light external to the HMD 1-100 from reaching the user's eyes, including by Figure 1B 1-108 is shown in the front perspective view of the HMD 1-100. In at least one example, the HMD 1-100 may also include a curtain 1-124 that shields the second opening 1-154 between the housing 1-150 and the rear display assembly 1-120a-b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.
[0066] Figure 1B and Figure 1C Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1D to 1F Any other examples of devices, features, components and parts shown and described herein. Figures 1D to 1F Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.
[0067] Figure 1D An exploded view of an example of an HMD 1-200 is shown that includes various parts or components that are separated according to modularization and selective coupling of these components. For example, the HMD 1-200 may include a band 1-216 that is selectively coupled to the first electronic strip 1-205a and the second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b can be removably coupled to the display unit 1-202.
[0068] In addition, the HMD 1-200 may include an optical seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218 that may be removably coupled to the display unit 1-202, for example, on a first component including a display screen and a second display component. The lens 1-218 may include a custom prescription lens configured to correct vision. As noted, in Figure 1D Each of the parts shown in the exploded view of and described above can be removably coupled, attached, reattached, and replaced to update parts or swap out parts for different users. For example, bands such as band 1-216, optical seals such as optical seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a-b can be swapped out depending on the user so that these parts are customized to fit and correspond to a single user of the HMD 1-200.
[0069] Figure 1D Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figure 1B , Figure 1C and Figure 1E to Figure 1FAny other examples of devices, features, components and parts shown and described herein. Figure 1B , Figure 1C and Figure 1E to Figure 1F Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1D Examples of devices, features, components, and parts are shown.
[0070] Figure 1E An exploded view of an example of a display unit 1-306 of an HMD is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0071] In at least one example, the display unit 1-306 may also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the position of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, with each display screen 1-322a-b having at least one motor such that the motor can translate the display screens 1-322a-b to match the interpupillary distance of the user's eyes.
[0072] In at least one example, the display unit 1-306 may include a dial or button 1-328 that is depressible relative to the frame 1-350 and accessible by a user external to the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 may be manipulated by a user to cause a motor of the motor assembly 1-362 to adjust the position of the display screens 1-322a-b.
[0073] Figure 1E Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1B to 1D and Figure 1F Any other examples of devices, features, components and parts shown and described herein. Figures 1B to 1D and Figure 1FAny of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1E Examples of devices, features, components, and parts are shown.
[0074] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is shown. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the position of the first display subassembly 1-420a and the second display subassembly 1-420b of the rear display assembly 1-421, including the first corresponding display screen and the second corresponding display screen for interpupillary adjustment, as described above.
[0075] Figure 1F The various parts, systems and assemblies shown in exploded views herein are referenced Figures 1B to 1E and subsequent figures referenced in this disclosure are described in more detail. Figure 1F The display unit 1-406 shown can be used with Figures 1B to 1E The fixing mechanism shown is assembled and integrated, and the fixing mechanism includes electronic strips, belts and other components including optical seals, connecting components, etc.
[0076] Figure 1F Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1B to 1E Any other examples of devices, features, components and parts shown and described herein. Figures 1B to 1E Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1F Examples of devices, features, components, and parts are shown.
[0077] Figure 1G A perspective exploded view of a front cover assembly 3-100 of an HMD device described herein is shown, for example Figure 1G The front cover assembly 3-1 of the HMD 3-100 shown, or any other HMD device shown and described herein. Figure 1GThe illustrated front cover assembly 3-100 may include a transparent or translucent cover 3-102, a shield 3-104 (or "canopy"), an adhesive layer 3-106, a display assembly 3-108 including a lenticular lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may secure the various components of the front cover assembly 3-100 to a frame or base of the HMD device.
[0078] In at least one example, Figure 1G As shown, the transparent cover 3-102, the shield 3-104 and the display assembly 3-108 including the lenticular lens array 3-110 can be bent to adapt to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 can be bent in two or three dimensions, for example, vertically bent in the Z direction inside and outside the ZX plane, and horizontally bent in the X direction inside and outside the ZX plane. In at least one example, the display assembly 3-108 may include a lenticular lens array 3-110 and a display panel having pixels that are configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., horizontally) to adapt to the curvature of the user's face from one side of the face (e.g., the left side) to the other side (e.g., the right side). In at least one example, each layer or component of the display assembly 3-108 (which will be shown in subsequent figures and described in more detail, but which may include the lenticular lens array 3-110 and the display layer) may be curved similarly or concentrically in the horizontal direction to accommodate the curvature of the user's face.
[0079] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back of the shield 3-104. When the HMD device is worn, the rear surface may be the surface of the shield 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104 opposite to the rear surface. In at least one example, one or more opaque portions of the shield 3-104 may include a peripheral portion that visually hides any components surrounding the outer periphery of the display screen of the display assembly 3-108. In this way, the opaque portion of the shield hides any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.
[0080] In at least one example, the shield 3-104 may define one or more aperture transparent portions 3-120 through which the sensor may send and receive signals. In one example, the portion 3-120 is a hole through which the sensor may extend or send and receive signals. In one example, the portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portion of the shield through which the sensor may send and receive signals through the shield and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0081] Figure 1G Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Figure 1G Examples of devices, features, components, and parts are shown.
[0082] Figure 1H An exploded view of an example of an HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102 including one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be fixed / secured.
[0083] Fig. 1I A portion of an HMD device 6-100 is shown including a front transparent cover 6-104 and a sensor system 6-102. The sensor system 6-102 may include a plurality of different sensors, emitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is shown in front of the sensor system 6-102 to illustrate the relative positions of the various sensors and emitters and the orientation of each sensor / emitter of the system 6-102. As referred to herein, "lateral," "sideways," "lateral," "horizontal," and other similar terms refer to aspects of the present invention. Figure 1J The x-axis is used to indicate the orientation or direction. Terms such as "vertical", "upward", "downward" and the like refer to the orientation or direction of the Figure 1J The orientation or direction is indicated by the Z-axis shown. Terms such as "forward," "rearward," "forward," "rearward" and the like refer to Figure 1J The Y-axis shown indicates the orientation or direction.
[0084] In at least one example, a transparent cover 6-104 may define a front exterior surface of the HMD device 6-100, and a sensor system 6-102 including various sensors and components thereof may be disposed in the Y axis / direction behind the cover 6-104. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted thereby.
[0085] As described elsewhere herein, the HMD device 6-100 may include one or more controllers including processors for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. In addition, as will be shown in more detail below with reference to other figures, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Fig. 1I Various structural frame members, brackets, etc. of the HMD device 6-100 are not shown. For clarity, Fig. 1I Components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.
[0086] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. The instructions may include or cause the processor to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein over time as the initial position, angle, or orientation of the camera is bumped or deformed due to an accidental drop event or other event.
[0087] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. The system 6-102 may include two scene cameras 6-102, respectively disposed on either side of the nose bridge or arch structure of the HMD device 6-100, such that each of the two cameras 6-106 roughly corresponds to the position of the left eye and the right eye of the user behind the cover 6-103. In at least one example, the scene camera 6-106 is generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene camera is a color camera and provides images and content for MR video pass-through to a display screen facing the user's eyes when the HMD device 6-100 is used. The scene camera 6-106 can also be used for environment and object reconstruction.
[0088] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 pointing generally forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction and hand and body tracking of a user. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centrally disposed along the width of the HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be disposed on an adapting structure above a central nose bridge or above a nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction and hand and body tracking. In at least one example, the second depth sensor may include a LIDAR sensor.
[0089] In at least one example, the sensor system 6-102 may include a depth projector 6-112 that is generally forward facing to project electromagnetic waves (e.g., in a predetermined pattern of light dots) into or within the field of view of a user and / or scene camera 6-106, or into or within a field of view that includes and exceeds the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light dots that are reflected from an object and returned to the depth sensors described above, including the depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 can be used for environment and object reconstruction and hand and body tracking.
[0090] In at least one example, the sensor system 6-102 may include downward facing cameras 6-114 whose fields of view are generally pointed downward on the Z axis relative to the HMD device 6-100. In at least one example, the downward cameras 6-114 may be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on a forward facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward cameras 6-114 may be used to capture facial expressions and movements of a user's face below the HMD device 6-100, including cheeks, mouth, and chin.
[0091] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on a forward-facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of a user's face beneath the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. Used for hand and body tracking, headset tracking, and facial avatars
[0092] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right side views in an X-axis or direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, headset tracking, and facial avatar detection and re-creation.
[0093] In at least one example, the sensor system 6-102 may include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of the user's eyes during and / or prior to use. In at least one example, the eye / gaze tracking sensors may include a nose-eye camera 6-120 that is disposed on either side of the user's nose and adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors may also include a bottom eye camera 6-122 disposed below the respective user's eyes for capturing images of the eyes for facial avatar detection and creation, gaze tracking, and iris identification functions.
[0094] In at least one example, the sensor system 6-102 may include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 may detect the overhead light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 may include a light emitting diode and may be particularly useful in low light environments for illuminating the user's hands and other objects in low light for detection by the infrared sensors of the sensor system 6-102.
[0095] In at least one example, multiple sensors (including the scene camera 6-106, the downward camera 6-114, the jaw camera 6-116, the side camera 6-118, the depth projector 6-112, and the depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination to better perform hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the above-described and Fig. 1I The downward camera 6-114, the jaw camera 6-116, and the side camera 6-118 shown in the figure can be wide-angle cameras capable of operating in the visible and infrared spectrum. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black and white light detection to simplify image processing and gain sensitivity.
[0096] Fig. 1I Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figure 1J to Figure 1L Any other examples of devices, features, components and parts shown and described herein. Figure 1J to Figure 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Fig. 1I Examples of devices, features, components, and parts are shown.
[0097] Figure 1J A lower perspective view of an example of an HMD 6-200 including a cover or shroud 6-204 secured to a frame 6-230 is shown. In at least one example, sensors 6-203 of a sensor system 6-202 may be disposed around the perimeter of the HMD 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of a display area or region 6-232 so as to not obstruct viewing of displayed light. In at least one example, the sensors may be disposed behind the shroud 6-204 and aligned with a transparent portion of the shroud, thereby allowing the sensor and projector to allow light to pass back and forth through the shroud 6-204. In at least one example, opaque ink or other opaque material or film / layer may be disposed on the shroud 6-204 around the display area 6-232 to hide components of the HMD 6-200 outside of the display area 6-232 rather than the transparent portion defined by the opaque portion through which the sensor and projector send and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass from the display (eg, within the display area 6-232), but does not allow light to pass radially outward from the display area around the display and the perimeter of the shield 6-204.
[0098] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 may send and receive signals. In the example shown, the sensor 6-203 of the sensor system 6-202 sends and receives signals through the shield 6-204, or more specifically, sends and receives signals through the transparent area 6-209 of (or defined by) the opaque portion 6-207 of the shield 6-204, which may include a sensor 6-203 having a plurality of transparent regions 6-209, and a plurality of transparent regions 6-209 defined by the opaque portion 6-207 of the shield 6-204. Fig. 1I The same or similar sensors as those shown in the example of FIG. 6-108, such as depth sensors 6-108 and 6-110, depth projector 6-112, first and second scene cameras 6-106, first and second downward cameras 6-114, first and second side cameras 6-118, and first and second infrared illuminators 6-124. These sensors are also Figure 1K and Figure 1L Other sensors, sensor types, number of sensors, and their relative positions may be included in one or more other examples of the HMD.
[0099] Figure 1J Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Fig. 1I and Figure 1K to Figure 1L Any other examples of devices, features, components and parts shown and described herein. Fig. 1I and Figure 1K to Figure 1L Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1J Examples of devices, features, components, and parts are shown.
[0100] Figure 1K A front view of a portion of an example of an HMD device 6-300 is shown, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shield in order to illustrate the brackets 6-336, 6-338. Figure 1J The illustrated shield 6-204 includes an opaque portion 6-207 that would visually cover / block viewing of anything outside (e.g., radially / peripherally outside) the display / display area 6-334, including the sensor 6-303 and the bracket 6-338.
[0101] In at least one example, the various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on angles relative to each other. For example, the tolerance on the mounting angles between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. In order to achieve and maintain such tight tolerances, in one example, the scene camera 6-306 may be mounted to the bracket 6-338 instead of the shield. The bracket may include a cantilever on which the scene camera 6-306 and other sensors of the sensor system 6-302 may be mounted to maintain position and orientation in the event of a drop by a user causing any deformation of the other brackets 6-226, the housing 6-330, and / or the shield.
[0102] Figure 1K Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1I to 1J and Figure 1L Any other examples of devices, features, components and parts shown and described herein. Figures 1I to 1J and Figure 1L Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1K Examples of devices, features, components, and parts are shown.
[0103] Figure 1L A bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is shown. The sensor system 6-402 may be similar to other sensor systems described above and elsewhere herein, including with reference to Figures 1I to 1K As described. In at least one example, the jaw camera 6-416 may face downward to capture images of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the frame or housing 6-430 as shown. The frame or housing 6-430 may include one or more holes / openings 6-415 through which the jaw camera 6-416 can send and receive signals.
[0104] Figure 1L Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figures 1I to 1K Any other examples of devices, features, components and parts shown and described herein. Figures 1I to 1KAny of the features, components and / or parts shown and described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1L Examples of devices, features, components, and parts are shown.
[0105] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is shown, the IPD adjustment system comprising first and second optical modules 11.1.1-104a-b slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 may be coupled to a bracket 11.1.1-112 and include a button 11.1.1-114 in electrical communication with the motors 11.1.1-110a-b. In at least one example, the button 11.1.1-114 may be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuit components to cause the first and second motors 11.1.1-110a-b to activate and respectively cause the first and second optical modules 11.1.1-104a-b to change position relative to each other.
[0106] In at least one example, the first and second optical modules 11.1.1-104a-b may include respective display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user may manipulate (e.g., press and / or rotate) the button 11.1.1-114 to activate position adjustment of the optical modules 11.1.1-104a-b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a-b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD so that the optical modules 11.1.1-104a-b may be adjusted to match the IPD.
[0107] In one example, a user may manipulate the button 11.1.1-114 to cause automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, a user may manipulate the button 11.1.1-114 to cause manual adjustment such that the optical modules 11.1.1-104a-b move farther or closer (e.g., as the user rotates the button 11.1.1-114 one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and power for moving the optical modules 11.1.1-104a-b via the motors 11.1.1-110a-b is provided by a power source. In one example, adjustment and movement of the optical modules 11.1.1-104a-b via the manipulation button 11.1.1-114 is mechanically actuated via the movement button 11.1.1-114.
[0108] Figure 1M Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components, and parts shown in any other figures and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to any other figures shown and described herein (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components, and parts shown in any other figures shown and described herein. Figure 1M Examples of devices, features, components, and parts are shown.
[0109] Figure 1N A front perspective view of a portion of an HMD 11.1.2-100 is shown, including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 defining a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. Figure 1N 106a-b may be blocked by one or more other components of the HMD 11.1.2-100 coupled to the internal frame 11.1.2-104 and / or the external frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the internal frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the internal frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.
[0110] The mounting bracket 11.1.2-108 may include a middle or center portion 11.1.2-109 coupled to the internal frame 11.1.2-104. In some examples, the middle or center portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the middle / center portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm extending away from the middle portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 extending away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 coupled to the internal frame 11.1.2-104.
[0111] like Figure 1N As shown, the external frame 11.1.2-102 may define a curved geometry on its underside to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry may be referred to as a nose bridge 11.1.2-111 and is centrally located on the underside of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the internal frame 11.1.2-104 between the holes 11.1.2-106a-b so that the cantilevers 11.1.2-112, 11.1.2-114 extend downwardly and laterally outward away from the middle portion 11.1.2-109 to complement the nose bridge 11.1.2-111 geometry of the external frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose, as described above. The geometry of the nose bridge 11.1.2-111 adapts to the nose as the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.
[0112] The first cantilever arm 11.1.2-112 may extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever arm 11.1.2-114 may extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite to the first direction. The first cantilever arm 11.1.2-112 and the second cantilever arm 11.1.2-114 are referred to as "cantilever" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 includes a free distal end 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 depend from the middle portion 11.1.2-109, which may be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are unattached.
[0113] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to a mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f may include various types of sensors, including cameras, IR sensors, and the like. In some examples, one or more of the sensors 11.1.2-110a-f may be used for object recognition in three-dimensional space, such that maintaining accurate relative positions of two or more of the plurality of sensors 11.1.2-110a-f is important. The cantilever nature of the mounting bracket 11.1.2-108 may protect the sensors 11.1.2-110a-f from damage and change of position in the event of an accidental drop by a user. Because the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, stresses and deformations of the inner frame and / or outer frame 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevered arms 11.1.2-112, 11.1.2-114 and therefore do not affect the relative positions of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.
[0114] Figure 1NAny of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included alone or in any combination in any other example of a device, feature, component described herein. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included alone or in any combination in any other example of a device, feature, component described herein. Figure 1N Examples of devices, features, components, and parts are shown.
[0115] Fig.1O An example of an optical module 11.3.2-100 for use in an electronic device such as an HMD, including an HMD device as described herein, is shown. As shown in one or more other examples described herein, the optical module 11.3.2-100 can be one of two optical modules within the HMD, where each optical module is aligned to project light toward an eye of a user. In this way, a first optical module can project light to a first eye of a user via a display screen, and a second optical module of the same device can project light to a second eye of a user via another display screen.
[0116] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the eyes of a user when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0117] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The camera 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the camera 11.3.2-106 is configured to capture one or more images of a user's eye during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the camera 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light emitting diodes (LEDs) or other lights configured to project light toward the eyes of the user when the HMD is worn. The individual lights 11.3.2-110 in the light strip 11.3.2-108 may be spaced around the light strip 11.3.2-108 and thus evenly or unevenly spaced around the display 11.3.2-104 at various locations on the light strip 11.3.2-108 and around the display 11.3.2-104.
[0118] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto the user's eyes. In one example, the camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.
[0119] As mentioned above, Fig.1O Each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (eg, second) optical module provided with the HMD to interact with (eg, project light and capture images) the user's other eye.
[0120] Fig.1O Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination. Figure 1P Any other examples of devices, features, components, and parts shown or otherwise described herein. Figure 1P Any of the features, components and / or parts shown or described herein (including arrangements and configurations thereof) may be included alone or in any combination. Fig.1OExamples of devices, features, components, and parts are shown.
[0121] Figure 1P A cross-sectional view of an example of an optical module 11.3.2-200 is shown, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first hole or channel 11.3.2-212 and a second hole or channel 11.3.2-214. The channels 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of the HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guides to secure the optical module 11.3.2-200 in place within the HMD.
[0122] In at least one example, the optical module 11.3.2-200 may also include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens that is removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light strip 11.3.2-208 and the one or more eye tracking cameras 11.3.2-206 such that the camera 11.3.2-206 is configured to capture images of the user's eyes through the lens 11.3.2-216, and the light strip 11.3.2-208 includes lights configured to project light into the user's eyes through the lens 11.3.2-216 during use.
[0123] Figure 1P Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included alone or in any combination in any other examples of devices, features, components, and parts described herein. Figure 1P Examples of devices, features, components, and parts are shown.
[0124] Figure 2is a block diagram of an example of a controller 110 according to some embodiments. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. To this end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., a universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0125] In some embodiments, the one or more communication buses 204 include circuits that interconnect and control communications between system components. In some embodiments, the one or more I / O devices 206 include at least one of a keyboard, a mouse, a touch pad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0126] The memory 220 includes a high-speed random access memory, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, the memory 220 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 optionally includes one or more storage devices located away from the one or more processing units 202. The memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 220 or the non-transitory computer-readable storage medium of the memory 220 stores the following programs, modules, and data structures or subsets thereof, including an optional operating system 230 and an XR experience module 240.
[0127] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., a single XR experience of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data sending unit 248.
[0128] In some embodiments, the data acquisition unit 241 is configured to obtain Figure 1A 120, and optionally acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, data acquisition unit 241 includes instructions and / or logic for instructions as well as heuristics and metadata for the heuristics.
[0129] In some embodiments, tracking unit 242 is configured to map scene 105 and track at least display generation component 120 relative to Figure 1A The tracking unit 242 may include instructions and / or logic for instructions and heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand and / or the position of one or more parts of the user's hand relative to the scene 105. Figure 1A The movement of the scene 105 relative to the display generation component 120 and / or relative to a coordinate system (the coordinate system is defined relative to the user's hand). Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the position or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hands)) or relative to the XR content displayed via the display generation component 120. Figure 5 The eye tracking unit 243 is described in more detail.
[0130] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripherals 195. To this end, in various embodiments, coordination unit 246 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.
[0131] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, position data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions as well as heuristics and metadata for the heuristics.
[0132] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are illustrated as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.
[0133] also, Figure 2 It serves more as a functional description of various features that may be present in a particular implementation, as opposed to a block diagram of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some embodiments, depends in part on the specific combination of hardware, software and / or firmware selected for a specific implementation.
[0134] Figure 31 is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. For this purpose, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal-facing and / or external-facing image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0135] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communications between various system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0136] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to the user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), microelectromechanical system (MEMS) and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffraction, reflection, polarization, holographic and other waveguide displays. For example, the display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 can present MR and VR content. In some embodiments, one or more XR displays 312 can present MR or VR content.
[0137] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye tracking camera). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to the user's hands and, optionally, at least a portion of the user's arms (and may be referred to as a hand tracking camera). In some embodiments, the one or more image sensors 314 are configured to face forward so as to acquire image data corresponding to a scene that the user would see in the absence of the display generating component 120 (e.g., an HMD) (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0138] The memory 320 includes a high-speed random access memory, such as a DRAM, SRAM, DDR RAM, or other random access solid-state memory device. In some embodiments, the memory 320 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 320 optionally includes one or more storage devices located away from the one or more processing units 302. The memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 320 or the non-transitory computer-readable storage medium of the memory 320 stores the following programs, modules, and data structures or a subset thereof, including an optional operating system 330 and an XR rendering module 340.
[0139] The operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, the XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0140] In some embodiments, the data acquisition unit 342 is configured to at least Figure 1A The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, positioning data, etc.). For the purposes described, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.
[0141] In some embodiments, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. For such purposes, in various embodiments, the XR rendering unit 344 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0142] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on the media content data. For such purposes, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0143] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, position data, etc.) to at least the controller 110, and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For such purposes, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0144] Although the data acquisition unit 342, the XR rendering unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., Figure 1A , but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in a separate computing device.
[0145] also, Figure 3 More as a functional description of various features that may be present in a particular embodiment, rather than a schematic diagram of the structures of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some embodiments, depends in part on the specific combination of hardware, software and / or firmware selected for a specific implementation.
[0146] Figure 4 is a schematic illustration of an example implementation of the hand tracking device 140. In some implementations, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the position / location of one or more parts of a user's hand, and / or one or more parts of a user's hand relative to Figure 1AThe hand tracking device 140 can be used to monitor movement of the scene 105 (e.g., relative to a portion of the physical environment surrounding the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hands). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to the head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0147] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least a human user's hand 406. The image sensor 404 captures hand images with sufficient resolution to enable the fingers and their corresponding positioning to be distinguished. The image sensor 404 typically captures images of other parts of the user's body, or may also capture images of all parts of the body, and may have zoom capabilities or a dedicated sensor with increased magnification to capture images of the hand with a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of the scene 105, or is used as an image sensor to capture the physical environment of the scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a manner that the field of view of the image sensor 404 or a portion thereof is used to define an interaction space, in which the hand movements captured by the image sensor are considered as input to the controller 110.
[0148] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D image data (and, in addition, possible color image data) to the controller 110, which extracts high-level information from the image data. The high-level information is typically provided to an application running on the controller via an application program interface (API), which drives the display generation component 120 accordingly. For example, a user can interact with the software running on the controller 110 by moving his hand 406 and changing his hand posture.
[0149] In some embodiments, the image sensor 404 projects a speckled pattern onto a scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the spots in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of points in the scene at a specific distance from the image sensor 404 relative to a predetermined reference plane. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis, so that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods based on a single or multiple cameras or other types of sensors, such as stereo imaging or time-of-flight measurement.
[0150] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or the processor in the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D position of the user's hand joints and finger tips.
[0151] The software can also analyze the trajectory of the hand and / or finger over multiple frames in the sequence to identify gestures. The pose estimation function described herein can be alternated with the motion tracking function so that the image block-based pose estimation is performed only once every two (or more) frames, and the tracking is used to find the changes in pose that occur on the remaining frames. The pose, motion, and gesture information is provided to the application running on the controller 110 via the above-mentioned API. The program can, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.
[0152] In some embodiments, gestures include air gestures. An air gesture is a gesture detected without the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) (or independent of an input element that is part of a device) and based on detected movement of a part of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body)).
[0153] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by movement of a user's fingers relative to other fingers or parts of the user's hand. In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture including a predetermined speed or rotation amount of a part of the user's body)).
[0154] In some embodiments where the input gesture is an in-air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in embodiments involving in-air gestures, for example, the input gesture is combined (e.g., simultaneously) with movement of the user's fingers and / or hand to detect attention (e.g., gaze) toward a user interface element to perform a pinch and / or tap input, as described below.
[0155] In some embodiments, an input gesture pointing to a user interface object is performed directly or indirectly with reference to the user interface object. For example, the user input is performed directly on the user interface object according to performing the input with the user's hand at a location corresponding to the location of the user interface object in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, when the user's attention (e.g., gaze) to the user interface object is detected, the input gesture is performed indirectly on the user interface object according to the location of the user's hand not being at the location corresponding to the location of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can guide the user's input to the user interface object by initiating a gesture at or near a location corresponding to the display location of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge of the option or the center portion of the option). For an indirect input gesture, the user can guide the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location that does not correspond to the display location of the user interface object).
[0156] In some embodiments, according to some embodiments, the input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch input and tap input for interacting with a virtual or mixed reality environment. For example, the pinch input and tap input described below are performed as air gestures.
[0157] In some embodiments, the pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes movement of two or more fingers of a hand to contact each other, that is, optionally followed by an immediate (e.g., within 0 to 1 second) interruption of contact with each other. A long pinch gesture as an air gesture includes movement of two or more fingers of a hand in contact with each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact with each other is detected. For example, a long pinch gesture includes the user maintaining a pinch gesture (e.g., in which two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected consecutively immediately (e.g., within a predefined time period) with each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).
[0158] In some embodiments, the pinch and drag gesture as an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., following) a drag input that changes the position of the user's hand from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input, and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers to contact each other and moves the same hand to a second position in the air using the drag gesture). In some embodiments, a pinch input is performed by a first hand of a user, and a drag input is performed by a second hand of the user (e.g., the second hand of the user moves from a first position to a second position in the air while the user continues a pinch input with the user's first hand. In some embodiments, an input gesture that is an air gesture includes input (e.g., pinch and / or tap input) performed using both hands of the user. For example, the input gesture includes two (e.g., or more) pinch inputs that are performed in conjunction with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) is performed using the first hand of the user, and in conjunction with the pinch input performed using the first hand, a second pinch input is performed using another hand (e.g., a second hand of the user's two hands).
[0159] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes movement of a user's finger toward the user interface element, movement of a user's hand toward the user interface element (optionally, extension of the user's finger toward the user interface element), downward movement of a user's finger (e.g., mimicking a mouse click motion or a tap on a touch screen), or other predefined movement of a user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture movement of the finger or hand, which is a movement of the finger or hand away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of the movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the acceleration direction of the movement of the finger or hand).
[0160] In some embodiments, it is determined that the user's attention is directed to a portion of the three-dimensional environment based on detection of a gaze directed to that portion of the three-dimensional environment (optionally, no other conditions are required). In some embodiments, it is determined that the user's attention is directed to a portion of the three-dimensional environment based on detection of a gaze directed to that portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., a dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, so that the device determines that the user's attention is directed to the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0161] In some embodiments, detection of a ready state configuration of a user or a portion of a user is detected by a computer system. Detection of a ready state configuration of a hand is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand. For example, the ready state of the hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grab gesture, or a pre-tap with one or more fingers extended and the palm facing away from the user), based on whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., toward an area in front of the user above the user's waist and below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface responds to attention (e.g., gaze) input.
[0162] In scenarios where input is described with reference to in-air gestures, it should be understood that similar gestures may be detected using a hardware input device attached to or held by one or more hands of a user, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units may be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used in place of the positioning and / or movement of the one or more hands in the corresponding in-air gesture. In scenarios where input is described with reference to in-air poses, it should be understood that similar poses may be detected using a hardware input device attached to or held by one or more hands of a user. User input may be detected using controls contained in a hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the position or change of position of parts of a hand and / or finger relative to each other, relative to the user's body and / or relative to the user's physical environment, and / or other hardware input device controls, wherein user input using the controls contained in the hardware input device is used in place of hand and / or finger gestures such as air taps or air pinches in corresponding air gestures. For example, a selection input described as being performed using an air tap or air pinch input may alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag may alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input following movement of the hardware input device (e.g., along with a hand associated with the hardware input device) through space). Similarly, two-handed input involving movement of the hands relative to each other may be performed using an air gesture and a hardware input device in the hand that is not performing the air gesture, two hardware input devices held in different hands, or two air gestures performed by different hands using various combinations of air gestures and / or inputs detected by one or more of the above hardware input devices.
[0163] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or may alternatively be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is also stored in a memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although in Figure 4Controller 110 is shown in FIG. 1 , but some or all of the processing functions of the controller may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other device associated with the image sensor 404, for example, as a separate unit from the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing functions of the image sensor 404 may also be integrated into a computer or other computerized device to be controlled by the sensor output.
[0164] Figure 4 Also included is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels with corresponding depth values. Pixels 412 corresponding to the hand 406 have been segmented from the background and wrist in the figure. The brightness of each pixel in the depth map 410 is inversely proportional to its depth value (i.e., the measured z distance from the image sensor 404), where gray shades become darker as depth increases. The controller 110 processes these depth values in order to identify and segment components of the image that have human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and motion from frame to frame in the depth map sequence.
[0165] Figure 4 Also schematically illustrated is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. Figure 4 In the hand skeleton 414, a hand background 416 that has been segmented from the original depth map is superimposed. In some embodiments, key feature points of the hand and optionally on a wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, palm center, end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points over multiple image frames to determine the gesture performed by the hand or the current state of the hand according to some embodiments.
[0166] Figure 5 The eye tracking device 130 ( Figure 1A ). In some embodiments, the eye tracking device 130 is composed of an eye tracking unit 243 ( Figure 2) controls to track the position and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as headphones, helmets, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head mounted device, and is optionally used in conjunction with a head mounted display generation component. In some embodiments, the eye tracking device 130 is not a head mounted device, and is optionally part of a non-head mounted display generation component.
[0167] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing a 3D virtual view to the user. For example, the head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture a video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and display a virtual object on the transparent or translucent display, and the user can directly view the physical environment through the transparent or translucent display. In some embodiments, the display generation component projects the virtual object into the physical environment. The virtual object may, for example, be projected on a physical surface or projected as a hologram so that an individual using the system observes the virtual object superimposed on the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.
[0168] like Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be pointed at the user's eyes to receive the IR or NIR light reflected directly from the eyes by the light source, or alternatively can be pointed at "hot" mirrors located between the user's eyes and the display panel, which reflect the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, the user's two eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by the corresponding eye tracking camera and illumination source.
[0169] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a specific operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, thermal mirror (if present), eye lens, and display screen. The device-specific calibration process can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process may include an estimate of the eye parameters of a specific user, such as pupil position, foveal position, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's gaze point relative to the display.
[0170] like Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system, which includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near infrared (NIR) camera) positioned on the side of the user's face on which eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 can be directed toward a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass) located between the user's eye 592 and a display 510 (e.g., a left display panel or a right display panel of a head-mounted display, or a display of a handheld device, a projector, etc.) (e.g., as Figure 5 ), or alternatively may be directed toward the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion of Figure 5 ).
[0171] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction in which the user is currently looking.
[0172] Several possible use cases for the user's current gaze direction are described below and are not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal area determined according to the user's current gaze direction than in the peripheral area. As another example, the controller may position or move virtual content in a view based at least in part on the user's current gaze direction. As another example, the controller may display specific virtual content in a view based at least in part on the user's current gaze direction. As another example use case in an AR application, the controller 110 may guide an external camera for capturing the physical environment of the XR experience to focus in the determined direction. The autofocus mechanism of the external camera may then focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses gaze tracking information to adjust the focus of the eye lens 520 so that the virtual object that the user is currently looking at has an appropriate degree of convergence to match the convergence of the user's eye 592. The controller 110 can use the gaze tracking information to guide the eye lens 520 to adjust the focus so that nearby objects that the user is looking at appear at the correct distance.
[0173] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)) mounted in a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light sources can be arranged in a ring or circle around each of the lenses, such as Figure 5 In some embodiments, for example, eight illumination sources 530 (eg, LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 may be used, and other arrangements and locations of illumination sources 530 may be used.
[0174] In some embodiments, the display 510 emits light in the visible range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the positions and angles of the eye tracking cameras 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0175] like Figure 5 Embodiments of the gaze tracking system illustrated in the drawings may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.
[0176] Figure 6 A flash-assisted gaze tracking pipeline according to some embodiments is illustrated. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., Figure 1A and Figure 5 The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and the flash in the current frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and the flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.
[0177] like Figure 6 As shown, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input to the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each group of captured images can be input to the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.
[0178] At 610, for the currently captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as indicated at 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0179] At 640, if advancing from element 610, the current frame is analyzed to track pupils and glints based in part on previous information from previous frames. At 640, if advancing from element 630, the tracking state is initialized based on the pupils and glints detected in the current frame. The processing results at element 640 are checked to verify that the results of the tracking or detection can be credible. For example, the results can be checked to determine whether the pupil and a sufficient number of glints for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are not likely to be credible, at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the results are credible, the method advances to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze point.
[0180] Figure 6 It is intended to be used as an example of an eye tracking technology that can be used for a particular implementation. As recognized by one of ordinary skill in the art, according to various embodiments, other eye tracking technologies currently existing or developed in the future can be used in place of or in combination with the flash-assisted eye tracking technology described herein in a computer system 101 for providing an XR experience to a user.
[0181] In some embodiments, the captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed on top of a representation of the real-world environment 602 .
[0182] Thus, the description herein describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table present in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of a computer system or passively displayed via a transparent or translucent display of a computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, wherein the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment so that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in a three-dimensional environment by placing virtual objects at corresponding locations in the three-dimensional environment that have corresponding locations in the real world to appear as if the virtual objects exist in the real world (e.g., a physical environment). For example, the computer system optionally displays a vase so that the vase appears as if a real vase is placed on top of a table in the physical environment. In some embodiments, a corresponding position in the three-dimensional environment has a corresponding position in the physical environment. Thus, when a computer system is described as displaying a virtual object at a corresponding position relative to a physical object (e.g., such as a position at or near a user's hand or at or near a physical table), the computer system displays the virtual object at a specific position in the three-dimensional environment so that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a position in the three-dimensional environment that corresponds to the position in the physical environment where the virtual object would be displayed if it were a real object at that specific position).
[0183] In some embodiments, real-world objects present in the physical environment that are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that only exist in the three-dimensional environment. For example, the three-dimensional environment may include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.
[0184] In a three-dimensional environment (e.g., a real environment, a virtual environment, or an environment comprising a mixture of real objects and virtual objects), an object is sometimes referred to as having depth or simulated depth, or an object is referred to as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension different from height or width. In some embodiments, depth is defined relative to a fixed coordinate set (e.g., where a room or object has a height, depth, and width defined relative to a fixed coordinate set). In some embodiments, depth is defined relative to the user's position or viewpoint, in which case the depth dimension varies based on the position and angle of the user's position and / or the user's viewpoint. In some embodiments where depth is defined relative to the user's position relative to the surface of the environment (e.g., the surface of the floor or ground of the environment), an object farther away from the user along a line extending parallel to the surface is considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder extending from the user's head toward the user's feet). In some embodiments where depth is defined relative to a user's viewpoint (e.g., relative to a direction of a point in space that determines which portion of an environment is visible via a head-mounted device or other display), objects that are farther away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's viewpoint and parallel to the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system with the origin of the viewpoint at the center of a sphere extending outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which applications and / or system content are displayed), where the user interface container has a height and / or width, and the depth is a dimension orthogonal to the height and / or width of the user interface container. In some embodiments, where the depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or is initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container is generally orthogonal or substantially orthogonal to a straight line extending from a user-based position (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments, where the depth is defined relative to the user interface container, the depth of an object relative to the user interface container refers to the position of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending away from the user or the user's viewpoint in different directions and / or from different starting points).In some embodiments, when depth is defined relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewpoint changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during an in-person collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content that includes the container). In some embodiments, for curved containers (e.g., including containers with curved surfaces or curved content areas), the depth dimension optionally extends into the surface of the curved container. In some cases, z separation (e.g., the separation of two objects in the depth dimension), z height (e.g., the distance of one object from another object in the depth dimension), z position (e.g., the position of an object in the depth dimension), z depth (e.g., the position of an object in the depth dimension), or simulated z dimension (e.g., depth used as a dimension of an object, a dimension of an environment, a direction in space, and / or a direction in simulated space) is used to refer to the concept of depth as described above.
[0185] In some embodiments, the user is optionally able to use one or both hands to interact with virtual objects in a three-dimensional environment as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of a computer system optionally capture one or more hands of a user and display representations of the user's hands in a three-dimensional environment (e.g., in a manner similar to displaying real-world objects in a three-dimensional environment as described above), or in some embodiments, due to the transparency / translucency of a portion of a display generation component that is displaying a user interface, or due to a projection of a user interface onto a transparent / translucent surface or a projection of a user interface onto a user's eyes or into the field of view of a user's eyes, the user's hands can be seen via the display generation component, via the ability to see the physical environment through the user interface. Therefore, in some embodiments, the user's hands are displayed at corresponding locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment, which can interact with virtual objects in the three-dimensional environment as if these virtual objects were physical objects in the physical environment. In some embodiments, the computer system can update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0186] In some of the embodiments described below, the computer system is optionally capable of determining an "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment, for example, to determine whether the physical object is directly interacting with the virtual object (e.g., whether the hand is touching, grabbing, holding, etc. a virtual object or is within a threshold distance of the virtual object). For example, a hand that directly interacts with a virtual object optionally includes one or more of the following: a finger of a hand pressing a virtual button, a user's hand grabbing a virtual vase, a user's hand closing together and pinching / holding the user interface of an application, and two fingers that perform any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a specific location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a specific corresponding location in the three-dimensional environment (e.g., if the hand is a virtual hand instead of a physical hand, the location where the hand will be displayed in the three-dimensional environment). The location of the hand in the three-dimensional environment is optionally compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing the location in the physical world (e.g., instead of comparing the location in the three-dimensional environment). For example, when determining the distance between the one or more hands of the user and the virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., if the virtual object is a physical object instead of a virtual object, the location where the virtual object will be located in the physical world), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same technology is optionally used to determine the distance between any physical object and any virtual object. Therefore, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, a computer system optionally executes any of the techniques described above to map the position of the physical object to a three-dimensional environment and / or to map the position of the virtual object to the physical environment.
[0187] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed to, and / or where and what the physical stylus held by the user is pointed to. For example, if the user's gaze is directed to a particular location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed to the virtual object. Similarly, the computer system is optionally able to determine the direction in which the stylus is pointing in the physical environment based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines a corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment that the stylus is pointing to, and optionally determines that the stylus is pointing to the corresponding virtual location in the three-dimensional environment.
[0188] Similarly, the embodiments described herein may refer to the position of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the position of a computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Therefore, in some embodiments, the position of the computer system is used as a proxy for the position of the user. In some embodiments, the position of the computer system and / or the user in the physical environment corresponds to the corresponding position in the three-dimensional environment. For example, the position of the computer system will be a position in the physical environment (and its corresponding position in the three-dimensional environment), and if the user stands at this position, facing the corresponding part of the physical environment visible via the display generation component, the user will see from this position in the physical environment in the same positioning, orientation and / or size (e.g., in an absolute sense and / or relative to each other) of the objects displayed in the three-dimensional environment by the display generation component of the computer system or visible in the three-dimensional environment via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed at the same locations in the physical environment as the virtual objects are located in the three-dimensional environment, and physical objects that have the same size and orientation in the physical environment as they do in the three-dimensional environment), then the position of the computer system and / or user is the position from which the user would see the positions of the virtual objects in the physical environment at the same positions, orientations, and / or sizes (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation components of the computer system.
[0189] In the present disclosure, various input methods are described with respect to interaction with a computer system. When an input device or input method is used to provide an example, and another input device or input method is used to provide another example, it should be understood that each example is compatible with the input device or input method described with respect to another example and optionally utilizes the input device or input method. Similarly, various output methods are described with respect to interaction with a computer system. When an output device or output method is used to provide an example, and another output device or output method is used to provide another example, it should be understood that each example is compatible with the output device or output method described with respect to another example and optionally utilizes the output device or output method. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment by a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with the method described with respect to another example and optionally utilizes these methods. Therefore, the present disclosure discloses an embodiment as a combination of features of multiple examples, without exhaustively listing all features of the embodiment in the description of each example embodiment.
[0190] User interface and associated processes
[0191] Attention is now turned to embodiments of a user interface ("UI") and associated processes that may be implemented on a computer system (such as a portable multifunction device or a head mounted device) having display generating components, one or more input devices, and (optionally) one or more cameras.
[0192] FIG. 7A to FIG. 7D An example of a computer system that displays virtual content showing areas of possible interaction and displays immersive virtual content according to some embodiments is shown.
[0193] Fig. 7A The computer system 101 is shown displaying a three-dimensional environment 702 via a display generation component (e.g., the display generation component 120 of FIG. 1 ) from the viewpoint of a user 701 shown in a top view (e.g., facing the back wall of the physical environment in which the computer system 101 is located). Figure 6 As described above, the computer system 101 optionally includes a display generation component (e.g., a touch screen) and a plurality of image sensors (e.g., Figure 3The image sensor 314 may include one or more of: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of a user or a portion of a user (e.g., one or more hands of a user) when the user interacts with the computer system 101. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display that includes a display generation component that displays a user interface or a three-dimensional environment to the user, and sensors that detect movement of the physical environment and / or the user's hands (such as movement interpreted by the computer system as gestures such as air gestures) (e.g., external sensors facing outward from the user), and / or sensors that detect the user's gaze (e.g., internal sensors facing inward toward the user's face).
[0194] like Fig. 7A As shown, the computer system 101 captures one or more images of the physical environment (e.g., operating environment 100) surrounding the computer system 101, including one or more objects in the physical environment surrounding the computer system 101. In some embodiments, the computer system 101 displays a representation of the physical environment in a three-dimensional environment 702, or portions of the physical environment are visible via the display generation component 120 of the computer system 101. For example, the three-dimensional environment 702 includes portions of the left and right walls, ceiling, and floor in the physical environment of the user 701, and also includes a physical object 706 that is a physical block and a physical object 710 that is a table.
[0195] exist Fig. 7A , three-dimensional environment 702 includes virtual content, such as virtual content 708A, virtual content 708B, and virtual content 704. Such virtual content is, optionally, any element displayed by computer system 101 that is not included in the physical environment of computer system 101.
[0196] In some embodiments, the virtual content 704 is displayed as an overlay on a portion of the physical environment (e.g., an outline). In some embodiments, the virtual content 704 corresponds to an area of the three-dimensional environment 702 with which the computer system 101 anticipates possible user interaction when displaying the virtual environment or other virtual content associated with the virtual content 708A, as will be described later. For example, the virtual content 704 and / or portions of the physical environment optionally correspond to a user's "viewing area." For example, when displaying the virtual environment or other virtual content associated with the virtual content 708A, the computer system 101 optionally anticipates that the user will likely stand within (e.g., stand in) an area of the physical environment corresponding to where the virtual content 704 is located. In some embodiments, the virtual content 708 optionally corresponds to a representation corresponding to the virtual environment (e.g., an immersive visual experience and / or an application that provides an immersive visual experience). In some embodiments, the computer system 101 initiates display of the virtual content at an immersion level greater than an immersion threshold in response to detecting input including a request to display such virtual content, as described in reference to FIG. Figure 7B As further described. The immersion level is described in more detail with reference to method 800. Thus, when computer system 101 is displaying a virtual environment or other virtual content associated with virtual content 708A, virtual content 704 is optionally a user-visual indication of a corresponding portion of the physical environment that user 701 may perceive. For example, if there is a physical object such as physical object 706 that warrants the user's attention, virtual content 704 draws the user's focus toward physical object 706. For example, when computer system 101 is displaying virtual content associated with virtual content 708A, such as Figure 7C As shown, there may be a risk of the user colliding with physical object 706. Thus, in some embodiments, virtual content 704 enhances the user's perception of the relationship between their physical space prior to interacting with such virtual content.
[0197] In some embodiments, virtual content 704 is displayed without displaying virtual content 708A and / or virtual content 708B. In some embodiments, virtual content 708A and / or 708B are displayed without displaying virtual content 704. In some embodiments, the visual appearance of virtual content 704, 708A, and 708B is similar to Fig. 7ADifferent from what is shown. For example, each virtual content is optionally displayed with different borders, lighting effects, colors, saturation, hue, brightness, animation, shape and / or position than what is shown. In some embodiments, virtual content 708B corresponds to a simulated shadow projected by virtual content 708A in response to one or more simulated light sources, which are optionally positioned above virtual content 708A but are optionally invisible. For example, a first simulated light source perpendicular to the floor of the physical environment and positioned above virtual content 708A optionally projects a virtual shadow (e.g., virtual content 708B) centered below virtual content 708A (e.g., projected onto virtual content 704). In some embodiments, the simulated light source is displayed and / or located at different positions and / or angles relative to virtual content 708A, so that as a supplement or alternative to virtual content 708B, additional virtual shadows of varying shapes, positions and / or intensities are displayed (e.g., on virtual content 704). Additionally or alternatively, one or more simulated light sources also cause the virtual content 708A to be displayed with a specular lighting effect, simulating the visual effect of real-world light shining on at least a semi-reflective surface, such that bright areas or spots are displayed on the virtual content 708A, indicating the position of the light source oriented toward the virtual content 708A.
[0198] Fig.7A1 Shown with Fig. 7A 704. For example, user 701 is located in an area of his physical environment corresponding to content 704 (e.g., Fig.7A1 outside the area indicated by the dotted line in FIG. Figure 7B Modifications to environment 702 are shown in response to input from user 701. Figure 7B As shown, the input includes the user 701 shown in the top view of the environment 702 moving to a position in the area of the environment 702 corresponding to the virtual content 704. The movement of the user 701 from outside the area corresponding to the virtual content 704 to inside the area corresponding to the virtual content 704 also moves from FIG. 7A1 to FIG. 7B2 704 (eg, Figure 7B2) within the area indicated by the dashed line in . In some embodiments, feedback and / or prompts are displayed in response to such input. For example, virtual content 712 (e.g., a confirmation prompt) is optionally displayed to ensure that the user wants the virtual content to be displayed at an immersion level greater than a threshold immersion level. In response to input corresponding to a request to display virtual content at an immersion level greater than the immersion threshold, the computer system 101 optionally displays virtual content 712 associated with the display of the virtual content at an immersion level greater than the immersion threshold. For example, the virtual content 712 optionally includes corresponding information associated with the virtual content to be displayed at an immersion level greater than the immersion threshold. The corresponding information optionally informs the user of the computer system 101 that the virtual content will be displayed (e.g., "the virtual environment will be loaded"). In some embodiments, the corresponding information includes a name associated with the virtual content (e.g., the name of the application that provides the virtual content to be displayed at the immersion level and / or the name of an immersive visual experience such as a beach, a forest, and / or a campsite). The corresponding information optionally also includes a prompt to confirm that the user is aware of their physical environment. For example, the corresponding information optionally includes a selectable option 712-1, which is selectable (e.g., using a mouse and cursor click, an attention and air gesture, an actuation of a physical and / or virtual button, and / or another suitable selection input pointing to the selectable option) to confirm the user's intention to display the virtual content at the immersive level. In some embodiments, the corresponding information optionally includes a selectable option 712-2, which is selectable to provide confirmation of the user's intention as previously described, and also abandon the display of at least a portion of the corresponding information 712 in response to a later received request to display the virtual content at the immersive level. For example, after receiving a selection of the selectable option 712-2, the computer system 101 is optionally made aware that the user does not wish to see the virtual content 712 and / or the selectable options 712-1 and 712-2 in the future. Therefore, at a later time, the computer system 101 detects an input corresponding to a request to load virtual content at the immersive level, and abandons the display of some or all of such virtual content as previously described, and optionally continues to display the virtual content at the immersive level. Thus, virtual content 712 helps computer system 101 and user 701 confirm the intent to display the virtual content and optionally reduces the need to continue displaying virtual content 712 .
[0199] As described with reference to method 800, in some embodiments, as part of the input, the computer system 101 detects that the position of a corresponding portion of the user 701 corresponds to a corresponding portion of the physical environment (e.g., corresponding to an area of the virtual content 704), which is referred to herein as a viewing area. In some embodiments, the computer system 101 is optionally unaware of which particular portion of the user corresponds to the viewing area. For example, a first input comprising movement of a user's feet into the area and a second input comprising movement of a user's hands into the area are optionally processed similarly or identically, such that virtual content 712 is optionally displayed in response to the first input and / or the second input. In some embodiments, the computer system 101 detects input based on movement of one or more expected portions of the user moving into the area. For example, the computer system 101 optionally displays virtual content 712 in response to detecting two feet of the user entering the area, rather than in response to a single foot entering the area and / or not in response to the user's hand entering the area. Thus, as Figure 7B As shown, computer system 101 displays virtual content 712 in response to the user's foot entering a corresponding area of the physical environment corresponding to virtual content 704 .
[0200] In some embodiments, additional virtual content associated with virtual content 704 is displayed in response to input. For example, computer system 101 optionally displays one or more selectable options, such as gripper 714-1, gripper 714-2, and / or gripper 714-3. In some embodiments, computer system 101 detects input associated with virtual content 704 directed to gripper 714-1, gripper 714-2, and / or gripper 714-3, and modifies one or more dimensions of virtual content 704. For example, computer system 101 optionally detects the user's attention (e.g., gaze) directed to the corresponding selectable option 714 while detecting an air gesture of hand 703A. For example, the air gesture is optionally an air pinch gesture including contact of the index finger and thumb of hand 703A. In some embodiments, the input includes movement of hand 703A while maintaining the air pinch gesture. For example, while the mid-air pinch gesture is maintained, the computer system 101 detects movement of the hand and, based on the movement, modifies one or more dimensions of the virtual content 704. For example, as indicated by attention 715B, the computer system 101 detects movement of the hand 703A while maintaining the mid-air pinch gesture, and scales (e.g., lengthens) the virtual content 704 based on the movement of the hand 703A away from the user 701 and / or scales (e.g., shrinks) the virtual content 704 based on the movement of the hand 703A toward the user 701 parallel to a first dimension (e.g., depth) of the virtual content 704.
[0201] In some embodiments, the computer system 101 scales the virtual content 704 by a scaling amount in a first direction based on the magnitude of the component of the movement of the hand 703A parallel to the first direction, without regard to movement of the hand in a second direction different from the first direction. For example, as described with reference to the gripper 714-1, the computer system 101 optionally detects movement of the hand 703A away from the user and to the left of the user while maintaining the air pinch gesture and while the user's attention is directed to the gripper 714-1, and discards consideration of the magnitude of the movement in the left direction, and scales the virtual content 704 based only on the magnitude of the component of the movement toward or away from the user 701 (e.g., parallel to the depth of the virtual content 704). Similarly, with reference to the gripper 714-3, the computer system 101 optionally scales the virtual content 704 based on the magnitude of the leftward and / or rightward movement of the hand 703A, and discards consideration of the magnitude of the movement toward and / or away from the user 701. In some embodiments, the computer system 101 scales the virtual content 704 along multiple dimensions according to the movement in multiple directions. For example, with reference to the grabber 714-2, the computer system 101 optionally scales the virtual content 704 according to the movement magnitude of the hand 703A toward the user 701, away from the user, to the user's left, and / or to the user's right to scale the width and / or length of the virtual content 704. In some embodiments, the user's movement magnitude scales the virtual content 704 equally in multiple directions. For example, moving the hand 703A forward in a first direction by a first movement magnitude may optionally scale the virtual content 704 equally along the first dimension and the second dimension (e.g., its depth and width) by a first amount. Similarly, moving the hand 703A to the right by a first movement magnitude may optionally scale the virtual content 704 by a first amount along the first dimension and the second dimension.
[0202] In some embodiments, the computer system 101 optionally abandons display of the virtual content 712 based on the satisfaction of one or more criteria, as further described with reference to method 800. For example, the computer system 101 is optionally made aware that the user 701 has recently received input requesting display of virtual content at an immersion level greater than an immersion threshold, and therefore abandons display of the virtual content 712. Such a scenario is optionally beneficial when the user temporarily or mistakenly moves outside the boundaries of the virtual content 704, so that upon re-entering the boundaries, the computer system 101 optionally abandons redundantly prompting the user to confirm their intention to display the virtual content at the immersion level. In some embodiments, the virtual content 712 is displayed with corresponding opacity and / or other visual characteristics (e.g., brightness, color, border, and / or visual effects) so that the user does not mistakenly ignore the virtual content 712. For example, the virtual content 712 is optionally completely opaque and optionally displayed with a colored border.
[0203] In some embodiments, in response to the selection of selectable options 712-1 and / or 712-2, the computer system 101 initiates a process of evaluating the user's physical environment. The evaluation optionally includes a scan of the physical environment. In some embodiments, the evaluation is initiated before selecting selectable options 712-1 and / or 712-2, such as in response to an input that displays virtual content at an immersion level greater than an immersion threshold, in response to powering on the device, and / or in response to other user interactions with the computer system 101. In some embodiments, the computer system 101 displays a representation of the scan, such as a grid pattern overlaid on the scanned object. In some embodiments, the scan includes a viewing area and / or an area defined by virtual content 704 of the user's physical environment. In some embodiments, the computer system does not initiate displaying virtual content at an immersion level until such a scan is completed. In some embodiments, the scan includes most or all of the user's physical environment in front of the user's viewpoint, and a portion of the environment behind the user's viewpoint. In some embodiments, the scan includes one or more portions of the physical environment corresponding to the viewing area corresponding to the virtual content 704 (e.g., corresponding areas of the physical environment with which a user can interact) and / or one or more portions of the physical environment outside of the viewing area. In some embodiments, the computer system 101 optionally detects a selection of selectable options 712-1 and / or 712-2 and, in response to such a selection, initiates display of the virtual content at an immersion level greater than an immersion threshold, as described in further detail below, and / or ceases display of the virtual content 712.
[0204] Figure 7B1 Shown with Figure 7B It should be understood that unless otherwise indicated below, Figure 7B1 Shown with FIG. 7A to FIG. 7D Elements shown with the same reference numeral have one or more or all of the same characteristics. Figure 7B1 The computer system 101 includes a display generation component 120 (or the same as the display generation component 120). In some embodiments, the computer system 101 and the display generation component 120 each have FIG. 7A to FIG. 7D The computer system 101 shown in FIG. 1 and FIG. Figure 3 One or more characteristics of the display generation component 120 shown, and in some embodiments, FIG. 7A to FIG. 7D The computer system 101 and display generation component 120 shown have Figure 7B1 The computer system 101 and one or more characteristics of the display generation component 120 are shown.
[0205] exist Figure 7B1, the display generation component 120 includes one or more internal image sensors 314a oriented toward the user's face (e.g., reference Figure 5 The eye tracking camera 540 is described above. In some embodiments, the internal image sensor 314a is used for eye tracking (e.g., detecting the user's gaze). The internal image sensor 314a is optionally arranged on the left and right portions of the display generation component 120 to enable eye tracking of the user's left and right eyes. The display generation component 120 also includes external image sensors 314b and 314c facing outward from the user to detect and / or capture the physical environment and / or the movement of the user's hand. In some embodiments, the image sensors 314a, 314b, and 314c have reference FIG. 7A to FIG. 7D One or more characteristics of the image sensor 314 .
[0206] exist Figure 7B1 , the display generation component 120 is shown as displaying optionally corresponding to the reference FIG. 7A to FIG. 7D Content is described as content displayed and / or visible via display generation component 120. In some embodiments, the content is displayed by a single display (e.g., Figure 5 In some embodiments, the display generation component 120 includes two or more displays (e.g., a left display panel and a right display panel for the user's left eye and right eye, respectively, as shown in FIG. Figure 5 These displays have the ability to be combined (e.g., by the user's brain) to create Figure 7B1 The display output of the view showing the content.
[0207] The display generation component 120 has a corresponding Figure 7B1 The field of view of the content shown (e.g., the field of view captured by external image sensors 314b and 314c and / or visible to the user via display generation component 120, indicated by the dashed line in the top view). Because display generation component 120 is optionally a head-mounted device, the field of view of display generation component 120 is optionally the same as or similar to the user's field of view.
[0208] exist Figure 7B1 , the user is depicted performing an air pinch gesture (e.g., using hand 703A) to provide input to computer system 101, thereby providing user input directed to content displayed by computer system 101. This depiction is intended to be exemplary and not limiting; the user optionally uses different air gestures and / or uses the same techniques as described in reference to FIG. FIG. 7A to FIG. 7D Other forms of input are used to provide user input.
[0209] In some embodiments, the computer system 101 responds to the FIG. 7A to FIG. 7DThe user input.
[0210] exist Figure 7B1 In the example of , the user's hand is visible in the three-dimensional environment because it is within the field of view of the display generation component 120. That is, the user can optionally see any part of his or her own body within the field of view of the display generation component 120 in the three-dimensional environment. It should be understood that FIG. 7A to FIG. 7D One or more or all aspects of the present disclosure shown or described with reference thereto and / or described with reference to corresponding methods are optionally described with reference to Figure 7B1 Similar or analogous methods are implemented on the computer system 101 and the display generation unit 120 .
[0211] Figure 7C Shows the response to Figure 7B The input in selects option 712-1 to display the virtual content at an immersion level greater than the threshold immersion level. The virtual content 704 has been scaled based on the attention, selection, and request to scale the virtual content 704, such as Figure 7B Therefore, the virtual content 704 shown is relatively larger than Figure 7B Virtual content shown.
[0212] As described herein, displaying virtual content at an immersive level optionally includes any suitable manner of displaying virtual content that was not displayed prior to receiving input requesting display (e.g., so as to replace at least a portion of the visibility of the physical environment in the three-dimensional environment 702), and / or optionally includes modifying the visual characteristics of the virtual content, as described in more detail with reference to method 800. For example, virtual content 716 optionally corresponds to an immersive visual experience. Such an immersive visual experience optionally includes a displayed representation of a simulated real-world scene, such as a previously recorded video of a campsite. In some embodiments, the immersive visual experience optionally includes a depiction of a completely or nearly completely virtual environment (e.g., a simulated physical space). For example, as Figure 7C The virtual content 716 shown shows a virtual sky that is part of a virtual beach of the virtual environment. The virtual content 716 optionally includes additional virtual content, such as user interfaces of applications associated with the computer system 101, virtual avatars of users of other computer systems, virtual avatars that do not correspond to the user (e.g., non-user characters), virtual objects, and other suitable virtual content. In some embodiments, the display of the virtual content occurs gradually. For example, the computer system 101 optionally displays the virtual content from a corresponding portion of the user's field of view (e.g., the right side, the left side, the top, the center, the bottom, corresponding to other virtual content such as Fig. 7AThe computer system 101 may initiate display of the virtual content 716 by starting with a portion of the virtual content 708A at a previous location in the user's field of view, and / or a combination of one or more of these portions. For example, the computer system 101 may optionally initiate display of the virtual content 716 at an upper region of the user's field of view, and may optionally continue to display a portion of the virtual content 716 toward another corresponding portion (e.g., a lower region) of the user's field of view, such that the display of the virtual content 716 is displayed at a lower region of the user's field of view. Figure 7C The amount of virtual content 716 shown in the three-dimensional environment 702 is gradually revealed in the three-dimensional environment 702. Alternatively, the virtual content 716 is optionally displayed starting from the right side of the user's field of view and ending towards the left side of the user's field of view, or vice versa. Thus, in some embodiments, displaying virtual content at an immersion level greater than the immersion level optionally includes displaying virtual content that was not displayed when the input requesting display of the virtual content was received.
[0213] As previously described, in some embodiments, displaying virtual content at an immersion level greater than an immersion threshold optionally includes modifying the visual characteristics of the virtual content. For example, the computer system 101 optionally applies one or more visual effects, such as a blur effect, a feathering effect, and / or a modification of the color space of one or more corresponding portions of the virtual content 716 (e.g., a brightness and / or saturation slightly lower than the final brightness and / or saturation of the corresponding content). The one or more corresponding portions optionally include the most recently displayed portion of the virtual content 716. For example, when the virtual content 716 is loaded from the upper area of the user's field of view to the lower area of the user's field of view, the lowest corresponding portion of the virtual content is optionally blurred and / or feathered, thereby enhancing visual focus and reducing the sudden loading of such content. In some embodiments, after displaying the additional corresponding portion of the virtual content 716 at an immersion level greater than the immersion threshold, the computer system 101 modifies the display of the previously displayed corresponding portion of the virtual content. For example, the first corresponding portion previously located at the "bottom" of the displayed virtual content 716 is no longer located at the bottom, because the display of the second corresponding portion of the virtual content below the first corresponding portion continues, and therefore, the computer system 101 modifies the first corresponding portion to stop the display of the visual effect. For example, the first corresponding portion is optionally displayed with a certain saturation, translucency level and / or other visual effects. In some embodiments, the computer system 101 optionally displays the first portion of the virtual content 716 simultaneously or almost simultaneously, rather than gradually displaying the first portion along one or more directions (e.g., from left to right, from top to bottom and / or some combination thereof). For example, the computer system 101 optionally causes the entire first portion of the virtual content 716 to gradually fade in (e.g., increase opacity). In some embodiments, the fade-in includes a halo visual effect. The halo visual effect optionally includes increasing the opacity of the central portion of the first portion at a rate greater than increasing the opacity of the distal portion of the first portion of the virtual content 716.
[0214] In some embodiments, the computer system 101 at least temporarily continues to display a portion of the corresponding area of the user's physical environment while displaying the virtual content at an immersion level greater than the immersion threshold. For example, the computer system 101 optionally displays a first portion of the virtual content 716 so that the first portion occupies a majority of the user's field of view, but does not display a second portion of the virtual content at an immersion level greater than the immersion threshold for a period of time (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, 15 seconds, 25 seconds, 50 seconds, 100 seconds, or 500 seconds) and / or until the user provides explicit input (e.g., actuation of a physical or virtual button, input including a voice command, and / or an air gesture such as movement of a user's hand swiping downward toward the bottom of the user's field of view) to initiate display of the second portion at an immersion level greater than the immersion threshold. In some embodiments, the second portion includes a corresponding area corresponding to the virtual content 704. When the second portion of the virtual content is not displayed at an immersion level greater than the immersion threshold, the user optionally has visibility of physical objects in the user's physical environment (such as physical object 706), potential contours (such as raised portions of the floor and / or the curb of the sidewalk), and / or other elements of the user's environment. Such a visual configuration allows the user to study the details of their physical environment, clear areas where they may interact with obstacles, and / or move to corresponding portions of the area so that their movement and interaction with the physical environment (e.g., the floor surrounding the user) is unimpeded or at least known to the user. Thus, in some embodiments, the computer system optionally retains a display of the user's physical environment in which the user expects that he or she may interact, thereby improving the user's perception of their surroundings and reducing the likelihood that the user will encounter spatial conflicts and / or collisions when moving and interacting with virtual content.
[0215] In some embodiments, in response to displaying virtual content 716 at an immersion level greater than an immersion threshold, the computer system 101 modifies and / or stops displaying virtual content 704. For example, before displaying the virtual content at an immersion level greater than the immersion threshold, the computer system 101 optionally displays the virtual content 704 as an at least partially transparent ring or rectangle displayed covering the floor around the user. When initiating the display of the virtual content at an immersion level greater than the immersion threshold, the computer system 101 optionally stops the display of the transparent ring and / or rectangle and replaces the virtual content with the second virtual content. In some embodiments, the computer system 101 optionally does not stop the display of the virtual content 704, but modifies the display of the virtual content 704. The modified version of the virtual content 704 optionally has one or more characteristics of the display of the second virtual content, but it should be understood that the two embodiments (although similar) are optionally different. In some embodiments, the second virtual content is displayed with an animation. For example, the second virtual content optionally includes one or more simulated light sources that illuminate the representation of the user's floor. In some embodiments, the one or more simulated light sources include one or more concentric rings of such light emanating from the user's current location (e.g., from the user's feet). For example, the simulated light optionally begins at a point corresponding to a corresponding part of the user (such as their feet) and spreads outward over time toward an outer portion of the viewing area corresponding to the virtual content 704. In some embodiments, the rings additionally or alternatively include a display of lines emanating from the user and spreading out on the floor. In some embodiments, the second virtual content optionally includes pulses of simulated light throughout the viewing area corresponding to the virtual content 704. For example, the pulses optionally include rhythmic brightening and dimming of the viewing area. In some embodiments, the simulated light may be a rhythmic brightening and dimming of the viewing area. Figure 7B The second virtual content corresponds to a larger or smaller area of the representation of the user's physical environment than the virtual content 704 shown.
[0216] In some embodiments, the display of the second virtual content and / or the modification of the virtual content 704 (e.g., virtual content and effects applied at an area corresponding to the viewing area) occurs simultaneously, while the second (e.g., lower) portion of the virtual content 716 is not displayed at an immersion level greater than the immersion threshold, and the first portion of the virtual content 716 is displayed at an immersion level greater than the threshold. For example, the computer system optionally displays simulated lights scattered on the floor of the user environment, while the lower portion of the virtual content 716 is not displayed. In some embodiments, if the second virtual content is displayed with an animation, the computer system 101 stops the display of the animation after a threshold time period (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, 15 seconds, 25 seconds, 50 seconds, 100 seconds, or 500 seconds). In some embodiments, after the threshold time period, the computer system 101 displays a static visual indication, such as a ring indicating the boundary of the user's viewing area.
[0217] Fig.7D A representation of the user's physical environment corresponding to the viewing area is shown being replaced with virtual content. For example, after a first portion of virtual content 716 has been displayed for a period of time (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, 15 seconds, 25 seconds, 50 seconds, 100 seconds, or 500 seconds) and a second portion of virtual content 716 has not been displayed during the period of time, computer system 101 initiates display of the second portion of virtual content 716. In some embodiments, display of the second portion of virtual content 716 has a reference to Figure 7C Initiation of display of virtual content in the video (e.g., initiating display of a first portion of virtual content 716) describes one or more characteristics of the display of virtual content. For example, if computer system 101 initiates display of virtual content 716 from an upper area of a user's viewpoint, then after allowing the user to see the viewing area for a period of time, the computer system proceeds to initiate display of a second portion of virtual content 716 at an immersion level greater than the immersion threshold, starting from an upper area of the second portion and moving downward toward the floor of the environment until the second portion is fully displayed. Thus, after giving the user an opportunity to observe the viewing area and potentially clear the viewing area of objects, computer system 101 optionally continues to display a fully immersive visual experience. For example, the second portion of virtual content 716 optionally includes a bottom portion of the virtual environment, such as sand on a beach or water in the ocean. In some embodiments, replacing a representation of the user's environment with virtual content includes occluding physical objects within the environment. For example, in Fig.7D, physical object 706 is no longer visible because virtual content 716 is displayed at an immersion level greater than the threshold immersion level. Therefore, although physical object 706 still occupies physical space in the user environment, it no longer obstructs viewing of virtual content 716. In some embodiments, computer system 101 also optionally stops display of virtual content 704 while replacing the viewing area with a second portion of virtual content 716. In some embodiments, display of the first portion of virtual content and / or replacement of the second portion of virtual content includes fading in the first and / or second portions (e.g., a gradual increase in opacity), and in the case of fading in the second portion, also includes fading out of the viewing area. Figure 7C The virtual content 704 shown in the embodiment of the present invention is displayed in the user's viewing area. For example, the computer system 101 optionally increases the opacity of the second portion of the virtual content 716 at a first rate and / or reduces the opacity of the virtual content 704 at a second rate (optionally the same as or different from the first rate). In some embodiments, if virtual content not included in the virtual content 716 is displayed in the user's viewing area, the computer system also optionally replaces the display of the virtual content not included in the virtual content 716 with the corresponding virtual content in the virtual content 716. For example, before the computer system 101 initiates the display of the second portion of the virtual content 716, a virtual window corresponding to the application user interface is optionally displayed in the user's viewing area. However, in response to initiating the display of the second portion of the virtual content 716, in addition to replacing the representation of the user environment (e.g., viewing area), the computer system 101 optionally stops displaying and / or fades out the virtual window.
[0218] FIG. 8A to FIG. 8F 1 is a flowchart illustrating an exemplary method for displaying virtual content at a visual significance level greater than a threshold visual significance level according to some embodiments. In some embodiments, method 800 is performed at a computer system (e.g., computer system 101 in FIG. 1 , such as a tablet device, a smart phone, a wearable computer, or a head-mounted device), the computer system including a display generation component (e.g., FIG. 1 , Figure 3 and Figure 4 In some embodiments, method 800 is performed by storing in a non-transitory computer-readable storage medium and by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., Figure 1ASome operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0219] In some embodiments, in conjunction with one or more input devices and display generation components (such as Fig. 7A A computer system (such as a display generation component 120) that communicates with Fig. 7A Method 800 is performed at the computer system 101 shown. For example, a mobile device (e.g., a tablet computer, a smart phone, a media player, or a wearable device) or a computer or other electronic device. In some embodiments, the display generation component is a display (optionally a touch screen display) integrated with the electronic device, an external display such as a monitor, a projector, a television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users. In some embodiments, one or more input devices include a device capable of receiving user input (e.g., capturing user input and / or detecting user input) and sending information associated with the user input to the computer system. Examples of input devices include a touch screen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the computer system), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device, a hand motion sensor). In some embodiments, the computer system communicates with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touch screen, trackpad)). In some embodiments, the hand tracking device is a wearable device, such as a smart glove. In some embodiments, the hand tracking device is a handheld input device, such as a remote control or a stylus.
[0220] In some embodiments, the computer system detects (802a) via one or more input devices a first input corresponding to a request to display virtual content, such as a first input from user 701 such as Figure 7B and Figure 7B1 As shown Figure 7C As shown to show Figure 7C716, which will visually replace a portion of a representation of a physical environment in which a user of the computer system is located when using the computer system, such as corresponding to the location of user 701. For example, when a virtual reality (VR) or mixed reality (XR) environment (e.g., in some embodiments, the first three-dimensional environment is an extended reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment) is optionally displayed including a visual representation of an immersive visual experience (e.g., a virtual environment) (e.g., icons and / or shapes displayed on a physical floor), such as described with reference to method 1000, the computer system optionally detects movement of a user of the computer system and / or the user's viewpoint to a location in the physical environment corresponding to the visual representation (e.g., into the visual representation). In some embodiments, the request to display the virtual content includes actuation of a physical and / or virtual button. In some embodiments, the request to display the virtual content includes detecting a gesture and / or posture of a user's attention and / or a corresponding part of the user (e.g., a hand and / or finger of the user). In some embodiments, the first input includes a request to view an immersive virtual experience (e.g., a virtual environment), such as a mixed reality environment consisting primarily of virtual content. In some embodiments, the virtual content and / or virtual environment is a simulated three-dimensional environment that is optionally displayed in a three-dimensional environment in place of a representation of a physical environment (e.g., complete immersion) or optionally simultaneously with a representation of a physical environment (e.g., partial immersion). Some examples of virtual environments include lake environments, mountain environments, sunset scenes, sunrise scenes, night environments, grassland environments, and / or concert scenes. In some embodiments, the virtual environment is based on a real physical location, such as a museum and / or an aquarium. In some embodiments, the virtual environment is a location designed by an artist. Therefore, displaying a virtual environment in a three-dimensional environment optionally provides the user with a virtual experience as if the user were physically located in the virtual environment. In some embodiments, the first input is or includes a tap or hand air gesture in space, such as air pointing or air pinching at an icon or other selectable option in an augmented reality (AR) or virtual reality (VR) environment to launch and / or display a virtual environment; or using an interface controller in an AR or VR environment to provide input to select an icon or other selectable option to launch and / or display a virtual environment (such as the first virtual environment described later). In some embodiments, the first input includes a hand of a user of the computer system performing a pinch air gesture, in which the index finger and thumb of the user's hand are brought together and touched, while the user's attention is directed to the icon or selectable option. In some embodiments, the first input is an attention-only and / or gaze-only input (e.g., does not include input from one or more parts of the user other than those parts that provide attention input).
[0221] In some embodiments, in response to detecting a first input via one or more input devices and based on determining that the first input corresponds to a request to display virtual content at an immersion level greater than an immersion threshold (e.g., 10%, 30%, 50%, or 75% immersion) (802b), the computer system displays (802c) via a display generation component a visual indication corresponding to a corresponding area of the physical environment with which a user of the computer system is able to interact when the virtual content is displayed at an immersion level greater than the immersion threshold, such as a Figure 7C The virtual content 704 is shown, while a representation of the corresponding area of the physical environment is visible via a display generating component such as Figure 7CA portion of the environment 702 shown. For example, the computer system optionally detects a request to display virtual content, such as an XR and / or VR enhancement of the user's current environment. In some embodiments, the computer system is not currently displaying virtual content, or is displaying a first amount of virtual content (e.g., system user interface elements such as date, time, and computer system status), and determines that the first input corresponds to a request to initiate display of a second virtual content. In some embodiments, the first input includes a request to view an immersive XR or VR environment, such that the amount of virtual content visible and / or presented to the user of the computer system increases in response to the first input. In some embodiments, the computer system determines that the first input includes a request to display virtual content, such that the requested virtual content occupies a user's field of view greater than a threshold amount (e.g., 0.1 degrees, 1 degree, 3 degrees, 5 degrees, 10 degrees, 15 degrees, 30 degrees, 45 degrees, 90 degrees, or 120 degrees), while changing the user's orientation relative to the three-dimensional environment. In some embodiments, the computer system displays virtual content at an opacity level greater than an opacity threshold (e.g., 0.01%, 0.1%, 1%, 3%, 5%, 10%, 50%, or 90% opacity). In some embodiments, the immersion level includes an associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by a computer system obscures background content (e.g., content other than the virtual environment and / or virtual content) surrounding / behind the virtual environment, optionally including the number of items of background content displayed and / or displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generation component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed. In some embodiments, background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application), virtual objects that are not associated with or included in the virtual environment and / or virtual content (e.g., files generated by a computer system or representations of other users, etc.), and / or real objects (e.g., transparent objects representing real objects in the physical environment surrounding the user, which are visible so that they are displayed via a display generation component and / or visible via transparent or translucent components of the display generation component because the computer system does not block / impede their visibility through the display generation component).In some embodiments, at a low immersion level (e.g., a first immersion level), background, virtual and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with background content, which is optionally displayed at full brightness, color and / or translucency. In some embodiments, at a higher immersion level (e.g., a second immersion level higher than the first immersion level), background, virtual and / or real objects are displayed in an obstructed manner (e.g., dimmed, blurred or removed from the display). For example, a corresponding virtual environment with a high immersion level is optionally displayed without simultaneously displaying background content (e.g., in full screen or full immersion mode). As another example, a virtual environment displayed at a medium immersion level is optionally displayed simultaneously with background content that is dimmed, blurred or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary between background objects. For example, at a specific immersion level, one or more first background objects are optionally more visually de-emphasized (e.g., dimmed, blurred and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects stop displaying. As referred to herein, the visual prominence of virtual content optionally refers to displaying one or more portions of the virtual content with one or more visual characteristics such that the virtual content is optionally distinct and / or visible relative to the three-dimensionality perceived by the user of the computer system. In some embodiments, the visual prominence of the virtual content has one or more characteristics described with reference to displaying the virtual content at an immersion level greater than and / or less than an immersion threshold. For example, the computer system optionally displays the corresponding virtual content with one or more visual characteristics having corresponding values, such as virtual content displayed at an opacity and / or brightness level. For example, the opacity level is optionally 0% opacity (e.g., corresponding to invisible and / or fully translucent virtual content), 100% opacity (e.g., corresponding to fully visible and / or non-translucent virtual content) and / or other corresponding opacity percentages corresponding to a discrete and / or continuous range of opacity levels between 0% and 100%. For example, reducing the visual prominence of a portion of the virtual content optionally includes reducing the opacity of one or more portions of the portion of the virtual content to 0% opacity or to an opacity value below the current opacity value. For example, increasing the visual prominence of the portion of the virtual content optionally includes increasing the opacity of one or more portions of the portion of the virtual content to 100% or to an opacity value greater than a current opacity value.Similarly, reducing the visual prominence of the virtual content optionally includes reducing the brightness level of one or more portions of the virtual content (e.g., a fully darkened visual appearance toward a 0% brightness level or another brightness value lower than the current brightness level), and increasing the visual prominence of the virtual content optionally includes increasing the brightness level of one or more portions of the virtual content (e.g., a fully brightened visual appearance toward a 100% brightness level or another brightness value higher than the current brightness level). It should be understood that modifications to visual prominence optionally include additional or alternative visual characteristics (e.g., saturation, where increased saturation increases visual prominence, while decreased saturation decreases visual prominence; blur radius, where increased blur radius decreases visual prominence, while decreased blur radius increases visual prominence; contrast, where increased contrast values increase visual prominence, while decreased contrast values decrease visual prominence). Changing the visual prominence of an object may include changing multiple different visual characteristics (e.g., opacity, brightness, saturation, blur radius, and / or contrast). Additionally, when the visual prominence of a first object is increased relative to the visual prominence of a second object, the change in visual prominence may be by increasing the visual prominence of the first object, or decreasing the visual prominence of the second object, increasing the visual prominence of both objects in a manner that the first object is increased more than the second object, or decreasing the visual prominence of both objects in a manner that the first object is decreased less than the second object. It should also be understood that the foregoing description of the modification of visual prominence applies to the embodiments described herein.
[0222] In some embodiments, when displaying virtual content, such as Figure 7C Virtual content 716 is shown, such as optionally occluding background content (e.g., a representation of the user's real world environment, such as Figure 7C The computer system optionally displays a geometric indication of possible interaction areas, such as Figure 7C 704 is shown. In some embodiments, the visual indication is a circle, rectangle, and / or oval that overlays a corresponding area of the user's physical environment (e.g., the floor and / or the area above the floor). The visual indication is optionally presented to the user to indicate the area (e.g., the corresponding area) with which the user can interact (e.g., move around in it), thereby showing where potential spatial conflicts between the user and real-world objects exist. In some embodiments, the visual indication shows the boundaries of the corresponding area of the physical environment (e.g., the boundaries displayed overlaid on a representation of the corresponding area of the physical environment), such as the corresponding area of environment 702 corresponding to virtual content 704, as shown. Figure 7CIn some embodiments, the corresponding area of the physical environment is larger or smaller than the corresponding visual indication corresponding to the corresponding area of the physical environment. For example, the visual indication is optionally a geometric shape overlaid on a portion of the representation of the real-world floor, however, the corresponding area of the physical environment optionally corresponds to the entire floor and / or an area of the floor visible from the user's current viewpoint, such as Figure 7C The floor of environment 702 is shown. In some embodiments, prompts and visual indications (such as a cursor) that remove physical objects from corresponding areas of the physical environment Figure 7B and Figure 7B1 The virtual content 712 shown is displayed simultaneously.
[0223] In some embodiments, the computer system displays (802d) the virtual content via the display generation component at an immersion level greater than an immersion threshold, such as Figure 7C The virtual content 716 shown includes replacing at least a portion of a representation of a corresponding area of the physical environment with the virtual content after displaying a visual indication corresponding to the corresponding area of the physical environment with which a user of the computer system can interact when the virtual content is displayed at an immersion level greater than an immersion threshold, such as replacing the representation of the physical environment with the virtual content 716. Fig.7DAs shown. For example, the first three-dimensional environment corresponds to a mixed reality environment including an immersive virtual experience, which optionally includes one or more areas of virtual content. One or more areas of virtual content, for example, optionally constitute 90% of the mixed reality environment, and one or more areas not including virtual content constitute the remaining 10% of the mixed reality environment. In some embodiments, the immersive virtual experience includes a virtual environment that completely or almost completely occupies the user's field of view of the computer system. In some embodiments, when the user changes their physical position and / or orientation relative to the immersive virtual environment, the virtual content included in the virtual environment completely occupies the user's field of view, so that the user remains surrounded by the virtual content. For example, the computer system optionally displays a visual indication such as a circular or rectangular shape overlaid on a corresponding area of the user's physical environment (e.g., floor), indicating an area where the computer system expects to interact with the virtual content and / or an area where the user is optionally allowed to move while maintaining an immersive experience and / or an area where the computer system optionally allows one or more functions to be initiated. In some embodiments, the visual indication is initially displayed at a position relative to the user and the first three-dimensional environment (e.g., centered on the user's position and / or feet). In some embodiments, the visual indication is static. In some embodiments, the visual indication is animated for a period of time or continues to be animated. In some embodiments, the visual indication is partially transparent, so that the virtual or real world ground or floor is at least partially visible through the visual indication. In some embodiments, at least a portion of the visual indication includes a representation of the physical environment. In some embodiments, the visual indication is offset from the ground or floor so that the visual indication appears to be hovering or has a height (e.g., 1cm, 3cm, 5cm, 10cm, 100cm or 1000cm) relative to the ground. In some embodiments, the visual indication remains visible, and the corresponding area of the three-dimensional environment is visible from the user's viewpoint. In some embodiments, the computer system stops the display of the visual indication after a threshold time amount (e.g., 0.01 seconds, 0.1 seconds, 0.25 seconds, 0.5 seconds, 1 second, 2.5 seconds, 5 seconds or 10 seconds) and displays virtual content at the corresponding area. In some embodiments, if the first input corresponds to a request to display virtual content in the first three-dimensional environment with an immersion level less than an immersion threshold, the computer system does not display the visual indication. Temporarily displaying a visual indication corresponding to a corresponding area of the environment may increase user safety by deflecting the user toward the corresponding area in the user's physical environment and foreshadowing potential collisions with physical objects within the corresponding area.
[0224] In some embodiments, a corresponding area (such as a region corresponding to a physical environment) of the computer system that a user can interact with when the virtual content is displayed at an immersion level greater than the immersion threshold is displayed via the display generation component. Figure 7CThe visual indication corresponding to the viewing area of the virtual content 704 shown in the figure (e.g., as described with respect to step 802) includes (804a) based on determining that the user is located at a first location in the physical environment (804b), the visual indication corresponding to the corresponding area is a first visual indication (804c) corresponding to the first area of the physical environment, such as Fig.7D The position of the virtual content 704 shown. For example, the computer system optionally determines the position of the user relative to the physical environment, such as the position of a corresponding part of the user (e.g., the head of the user, the feet of the user, and / or the torso of the user) corresponds to a first position of the physical environment. In some embodiments, the computer system determines that the first position of the user corresponds to a corresponding area of the physical environment, referred to herein as a "physical viewing area". For example, the computer system optionally determines that the first position of the user at least partially intersects with the physical viewing area and / or is within the physical viewing area. In some embodiments, the visual indication referred to herein as a "viewing area" has one or more characteristics described in step 812. In some embodiments, the first area of the physical environment is defined relative to a portion of the user's viewpoint. For example, the computer system optionally determines that the first area of the physical environment corresponds to a portion of the user's field of view extending from the physical floor toward the physical ceiling or sky (e.g., 5%, 10%, 15%, 20%, 25%, 30%, 35%, or 40%). In some embodiments, the area of the physical environment corresponds to a portion (e.g., 5%, 10%, 15%, 20%, 25%, 30%, 35%, or 40%) of the physical floor relative to the user's viewpoint (e.g., centered on the user's physical location, such as at the user's feet). In some embodiments, the area corresponds to the area of the physical environment that is visible relative to the user's viewpoint (0.01 m 2 、0.05m 2 、0.1m 2 、0.5m 2 , 1m 2 、5m 2 、10m 2 ). In some embodiments, the first area of the physical environment has a world locked position.
[0225] In some embodiments, the computer system replaces (804d) a portion of a representation of a corresponding area of the physical environment with virtual content, including replacing at least a portion of a representation of a first area of the physical environment with the virtual content, such as by Fig.7D, as shown in virtual content 716 in the physical environment. For example, the computer system optionally at least partially or completely stops the display of the previously described physical viewing area and / or initiates the display of virtual content greater than the immersion threshold, as described with respect to step 802. In some embodiments, the physical viewing area is potentially visible (e.g., via passive visual transmission such as a transparent material sheet), however, the display of the virtual content obscures the visibility of the representation of the corresponding area. For example, if the user moves to a first position in the physical environment (e.g., has entered the viewing zone), the physical viewing area initially optionally occupies the lower area of the user's field of view, and the virtual content is visible.
[0226] In some embodiments, based on determining that the user is located at a second location in the physical environment that is different from the first location (804e), the visual indication corresponding to the corresponding area is a second visual indication corresponding to a second area of the physical environment that is different from the first area of the physical environment, such as a second visual indication displayed at a second location different from that shown. Figure 7C The virtual content 704 (804f) shown. For example, the visual indication is optionally displayed at a location different from the first location within the XR or VR environment, optionally corresponding to the corresponding portion of the user. In some embodiments, the computer system displays the visual indication at a corresponding area of the physical environment corresponding to the corresponding portion of the user.
[0227] In some embodiments, the computer system replaces a portion of the representation of the corresponding area of the physical environment with the virtual content, including replacing at least a portion of the representation of the second area of the physical environment with the virtual content, such as replacing physical object 706 with virtual content 716. Fig.7D As shown (804g) (e.g., the same or similar as described with respect to replacing the first area of the physical environment with virtual content). Replacing a portion of the representation of the corresponding area of the physical environment with virtual content based on determining that the user is located at the corresponding location in the physical environment provides a consistent visual experience regardless of changes in the corresponding location of the user, thereby reducing the likelihood that the user will incorrectly interact with the virtual content and / or reducing the need for input to redirect the virtual content relative to the corresponding location.
[0228] In some embodiments, visual indications corresponding to respective areas of the physical environment with which a user of the computer system can interact are displayed in association with the floor of the physical environment (806), such as Figure 7C704 of the virtual content shown. For example, the viewing area (e.g., visual indication) optionally corresponds to a portion of the floor of the user's physical environment so that the user is visually guided to the floor of the physical environment. In some embodiments, the portion of the floor is a circular, rectangular, or other shaped area of the floor of the physical environment, which is optionally centered on the user's feet. Displaying visual indications associated with corresponding areas of the physical environment provides information about potential spatial conflicts with the physical environment when simultaneously viewing virtual content, thereby improving user safety.
[0229] In some embodiments, the visual indication has a first shape, and wherein the visual indication is at least partially translucent, such as Figure 7C 704 (808) of the virtual content shown. For example, the visual indication is optionally a ring-shaped graphic overlaid on the representation of the user's physical floor, optionally showing the boundaries of the viewing area, and / or optionally displayed with a corresponding level of translucency (e.g., 5%, 10%, 15%, 20%, 25%, 35%, 45%, 60%, or 75% translucent). In some embodiments, the visual indication has one or more characteristics as described in method 1000. Displaying the visual indication with partial translucency reduces visual obstruction of the representation of the physical environment, thereby reducing the likelihood that the user will unexpectedly collide with portions of the physical environment.
[0230] In some embodiments, the first shape is elliptical and has a first corresponding diameter, and the visual indication includes a plurality of shapes, including the first shape and a second shape different from the first shape, wherein the second shape has a second diameter different from the first diameter, such as Figure 7C 810. In some embodiments, the visual indication includes a plurality of concentric shapes (e.g., rings). In some embodiments, the plurality of shapes are centered on the user's respective position. In some embodiments, the plurality of shapes are animated, similar to that described in step 812. For example, the plurality of concentric shapes optionally emanate from the user's respective position. Displaying a visual indication with a plurality of shapes draws the user's attention to the respective area of the physical environment, thereby reducing the likelihood that the user will unexpectedly collide with portions of the physical environment.
[0231] In some embodiments, displaying, via the display generation component, a visual indication corresponding to a respective area of the physical environment with which a user of the computer system can interact includes displaying a first position (such as a first position) corresponding to a respective portion of the user in the three-dimensional environment. Figure 7C The animation (such as the position of the virtual content 704 shown) extends to a visually indicated boundary of a second position in the three-dimensional environment different from the first position. Figure 7CIn some embodiments, the viewing area (e.g., a visual indication of a corresponding area of the physical environment with which the user may interact) has one or more characteristics of the animation described in step 802. For example, in response to displaying virtual content (such as a video clip) at an immersion level greater than an immersion threshold, the viewing area (e.g., a visual indication of a corresponding area of the physical environment with which the user may interact) has one or more characteristics of the animation described in step 802. Figure 7C The computer system optionally initially displays a visual indication having a first shaped boundary and a first size, such as a first input corresponding to a request for virtual content 716 as shown. Figure 7C The visual indication of a first boundary and a first size of the virtual content 704 shown (e.g., a relatively small circle centered on the user and / or the user's feet), optionally expanding over time to a second shaped boundary having a relatively larger size (e.g., a relatively larger circle) that is optionally similar to the first shape, such as Figure 7C The second size of the virtual content 704 shown. In some embodiments, the animation includes continuing to expand the border toward the outer end of the user's viewpoint or a maximum size defined by the computer system. In some embodiments, the animation also includes visual effects, such as glow effects, blur effects, modifications of translucency, lighting effects and / or modifications of brightness as described in step 814. In some embodiments, the border of the visual indication is continuously animated (e.g., expanded) until the outer end of the user's viewpoint and / or the maximum size is reached, such as Figure 7C Animation of virtual content 704 is shown. Animating visual cues draws the user's attention to corresponding areas of the physical environment, thereby reducing the likelihood that the user will unexpectedly collide with portions of the physical environment.
[0232] In some embodiments, the animation includes visual effects applied to surfaces of corresponding areas of a physical environment with which a user of a computer system can interact, such as applied to Figure 7C A visual effect (814) of the virtual content 704 shown. For example, the visual effect optionally has one or more characteristics as described in step 812, such as a simulated lighting effect optionally applied to a surface of the representation of the physical viewing area (e.g., a floor, a surface of a corresponding object located on top of the floor, and / or a physical wall). In some embodiments, the simulated lighting effect is based on one or more virtual light sources positioned and oriented toward a corresponding location of the physical environment. For example, the corresponding virtual light source is optionally visible or invisible and oriented perpendicular to the surface of the representation of the physical viewing area. Including the visual effect applied to the surface of the corresponding area of the physical environment will draw the user's attention to the corresponding objects included in the corresponding area and the outline of the corresponding area, thereby improving user safety.
[0233] In some embodiments, when a boundary of a visual indication corresponding to a corresponding area of the physical environment is animated via the display generation component, the computer system stops (816) animating the boundary of the visual indication based on determining that the visual indication has been animated for a time period greater than a threshold time period (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, 15 seconds, 30 seconds, or 60 seconds), such as Figure 7C Stopping the animation of the virtual content 704 shown. For example, after the visual indication has been animated for more than a threshold amount of time, the animation is optionally gradually or suddenly stopped. In some embodiments, after the threshold time has passed, the visual indication continues to be displayed with a default appearance (e.g., including one or more characteristics of the visual effect of the animation described in step 812). In some embodiments, the visual indication continues to be displayed after the animation stops with a visual appearance that matches the appearance of the visual effect when the animation stopped. Stopping the animation visually guides the user away from the corresponding area of the physical environment, thereby enhancing the focus on the displayed virtual content, and optionally indicates that a specific input that optionally does not initiate the execution of the function while the animation is ongoing is optionally operable to initiate the execution of the function.
[0234] In some embodiments, displaying the virtual content at an immersion level greater than the immersion threshold via the display generation component includes (818a) replacing a second representation (818b) of a second corresponding area of the representation of the physical environment corresponding to an upper area of the user's viewpoint with the first portion of the virtual content, such as Figure 7C The virtual content 716 shown to Fig.7D The virtual content shown. For example, the second corresponding area of the representation of the physical environment optionally includes a portion of the upper area of the user's field of view (e.g., 0.1 degrees, 1 degree, 3 degrees, 5 degrees, 10 degrees, 15 degrees, 30 degrees, 45 degrees, 90 degrees, or 120 degrees). In some embodiments, replacing the second representation of the second corresponding area includes reducing the visual significance of the second representation of the second corresponding area (e.g., stopping display and / or increasing the corresponding translucency). In some embodiments, the remaining (e.g., unreplaced) portion of the representation of the physical environment is maintained when the replacement occurs. For example, the replacement optionally includes an animation that gradually reduces the corresponding visual significance of the second corresponding area in a first direction (e.g., from the top of the user's field of view downward toward the bottom of the user's field of view) while maintaining the corresponding visual significance of the remaining part of the representation of the physical environment. Additionally or alternatively, when the replacement occurs, the visual significance of the first virtual content is optionally increased. For example, the animation includes gradually increasing the visual significance of the virtual content of the second representation that replaces the second corresponding area.
[0235] In some embodiments, after replacing the second representation of the second corresponding area, the computer system replaces a third representation of a third corresponding area of the representation of the physical environment with the second portion of the virtual content, the third corresponding area corresponding to a lower area of the user's viewpoint, lower than an upper area of the user's viewpoint, such as Figure 7C The virtual content 716 shown to Fig.7D The virtual content (818c) shown. For example, the replacement of the third representation of the third corresponding area of the representation of the physical environment is optionally initiated based on determining that one or more criteria are met, and the one or more criteria include the criteria that are met when a threshold amount of time (0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, or 15 seconds) has passed since the replacement of the second representation of the second corresponding area was initiated or completed. In some embodiments, the one or more criteria include the criteria that are met when a user input (e.g., actuation of a physical button, selection of a selectable affordance that stops the display of the third corresponding area, and / or input including movement detected within the corresponding area of the physical environment) is received. In some embodiments, the replacement of the third representation of the third corresponding area has one or more characteristics of the replacement of the second representation of the second corresponding area. Additionally or alternatively, when the replacement occurs, the visual prominence of the first virtual content is optionally increased. For example, the animation includes gradually increasing the visual prominence of the virtual content that replaces the second representation of the third corresponding area. In some embodiments, the animation described herein is included in an animation that continuously replaces the representation of the physical environment from the top of the user's field of view to the bottom of the user's field of view. Continuously replacing corresponding areas of the representation of the physical environment with corresponding virtual content visually directs the user's attention toward the lower area of the user's viewpoint, thereby reducing the possibility of spatial conflict between the user of the computer system and physical objects visible in the lower area of the user's viewpoint.
[0236] In some embodiments, in response to detecting a first input via one or more input devices, and based on determining that one or more criteria (including criteria that are satisfied when the first input corresponds to a request to display virtual content at an immersion level greater than an immersion threshold) are satisfied, the computer system displays (820) via a display generation component indicating that the virtual content is to be displayed at an immersion level greater than the immersion threshold, such as corresponding virtual content. Figure 7B and Figure 7B1The virtual content 712 shown. For example, the computer system optionally displays a virtual object including corresponding virtual content (e.g., text and / or a graphical icon) indicating that immersive virtual content will be loaded. In some embodiments, the corresponding virtual content displays a description of the virtual content. In some embodiments, the corresponding virtual content includes one or more selectable options associated with the display of the virtual object and / or its corresponding virtual content described in further detail below. In some embodiments, as described in step 824, the corresponding virtual content is displayed if one or more criteria are met. Displaying corresponding virtual content indicating that the virtual content will be displayed at an immersion level greater than the immersion threshold reduces the likelihood that the user will mistakenly initiate the display of the virtual content, thereby reducing the processing required to initiate such an erroneous display and preventing the need for input to eliminate the virtual content.
[0237] In some embodiments, the number of times the virtual content has been displayed at an immersion level greater than the immersion threshold (such as Figure 7C The one or more criteria (822) are satisfied regardless of the number of times the virtual content 716 shown has been displayed). For example, the virtual object described in step 820 is optionally displayed each time in response to the first input, regardless of the previous history of interactions associated with the virtual content (e.g., the number of times inputs similar to the first input have been received and / or the number of times the virtual content or other virtual content has been displayed at an immersion level greater than the immersion threshold). In some embodiments, as described in step 820, the virtual object includes one or more selectable options (e.g., including "confirm" and / or including "do not show again"). In some embodiments, in response to detecting an input selecting a corresponding enabling representation included in the corresponding virtual content, the computer system initiates display of the virtual content (e.g., the immersive visual experience described in step 802). In some embodiments, the one or more criteria include a criterion that is satisfied when the user has not previously selected the corresponding enabling representation included in the corresponding virtual content (e.g., "do not show again") (e.g., as described relative to step 824, the computer system optionally abandons display of the virtual object (e.g., the corresponding virtual content indicates that the virtual content will be displayed at an immersion level greater than the immersion threshold)). Displaying the corresponding virtual content independent of the number of times the virtual content has been displayed ensures that the user has a consistent expectation of what to display in response to the first input, thereby reducing the likelihood that the user will mistakenly direct the input to the virtual content.
[0238] In some embodiments, in response to detecting the first input via one or more input devices, and based on determining that one or more criteria are not met (e.g., as described in step 822), the computer system abandons (824) displaying the corresponding virtual content via the display generation component, such as Figure 7CThe aforementioned display of the virtual content 716 shown in step 820. For example, the one or more criteria include criteria that are not met when the user recently interacted with the corresponding virtual content (e.g., the corresponding virtual object) described in step 820. In some embodiments, the one or more criteria include criteria that are not met based on the recency of the interaction described in step 826. For example, if one or more criteria are not met, the computer system optionally abandons the display of the corresponding virtual content, such as the virtual object and / or virtual content. As another example, the computer system optionally determines that the user has recently provided input requesting that the virtual content be displayed at an immersion level greater than a threshold immersion level, and therefore abandons the display of the corresponding virtual content. Abandoning the display of the corresponding virtual content reduces the user input required to stop the display of the corresponding virtual content.
[0239] In some embodiments, the one or more criteria include criteria that are satisfied based on the recency of a previous interaction of the user of the computer system with the virtual content (826). For example, the computer system optionally forgoes display of the corresponding virtual content, such as Fig. 7A As shown, if the computer system detects that the user has recently interacted with the virtual content, such as displaying Figure 7B and Figure 7B1 If the computer system 101 performs a request for the virtual content 712 shown in FIG. 1 (e.g., initiates loading of the virtual content, releases the virtual content, and / or moves the corresponding virtual content into the virtual content and / or away from the virtual content), the computer system 101 abandons display of the virtual content 712, such as Figure 7B and Figure 7B1 In some embodiments, the one or more criteria include when the user has recently interacted with virtual content displayed at an immersion level greater than an immersion threshold (such as Figure 7C The corresponding criteria satisfied when the user interacts with the virtual content 716 shown in step 802 are similar or the same as described in step 802. In some embodiments, one or more criteria include when the user moves to the virtual content 716 within a threshold amount of time (e.g., 0.05 hours, 0.1 hours, 0.5 hours, 1 hour, 5 hours, 10 hours, 50 hours, 100 hours, or 500 hours) from receiving the first input (such as after detecting that the user 701 moves to the virtual content 716 shown in step 802). Figure 7B and Figure 7B1 Including criteria that are satisfied based on the recency of a previous interaction of a user of the computer system with the virtual content reduces the display of redundant corresponding virtual content that the user may not wish to view.
[0240] In some embodiments, one or more criteria include a criterion (828) that is satisfied based on the recency of detecting, via one or more input devices, a previously received corresponding input corresponding to a request to display corresponding virtual content in a corresponding area of the physical environment at an immersion level greater than an immersion threshold (such as the recency of detecting an input of hand 703A pointing to selectable option 712-1). For example, as described in steps 824-826. In some embodiments, the previously received corresponding input is the same as the first input described with respect to step 802. In some embodiments, the previously received corresponding input is a different input, such as an input that displays recently displayed virtual content (e.g., the immersive visual experience described in step 802). In some embodiments, the recency of detecting the corresponding input is based at least in part on the corresponding physical environment in which the user is located when the corresponding input is detected. For example, one or more criteria optionally include a criterion that is satisfied when the corresponding input is received while the user is in the same corresponding physical environment (e.g., the same room) as the current physical environment (e.g., the room). In some embodiments, one or more criteria include criteria that are not satisfied when the user is in a different corresponding physical environment (such as a first room) when a corresponding input different from the user's current physical environment (such as a second room different from the first room) is received. In some embodiments, one or more criteria include criteria that are satisfied when the current physical environment resembles the corresponding physical environment in which the user was when the corresponding input was received to a greater extent than a threshold amount (e.g., 5%, 10%, 15%, 25%, 35%, 50%, 65%, 75%, or 90%). For example, the current physical environment is optionally the first room, and the corresponding physical environment optionally corresponds to a doorway connecting the first room and a different second room. Including criteria that are satisfied based on the recency of detecting a previously received corresponding input corresponding to a request to display corresponding virtual content reduces the display of redundant corresponding virtual content due to the recency of the user providing such corresponding input when being in a similar physical environment.
[0241] In some embodiments, replacing at least a portion of the representation of the corresponding area of the physical environment (such as physical object 706) with virtual content after displaying a visual indication corresponding to the corresponding area of the physical environment with which a user of the computer system can interact includes maintaining display of at least a portion of the representation of the corresponding area of the physical environment, such as Fig.7DVirtual content 716 (830) is shown. For example, the computer system optionally maintains visibility of at least a portion of the physical viewing area and optionally replaces different portions of the physical viewing area with virtual content (e.g., part of an immersive visual experience), similar to that described in step 802. In some embodiments, as part of maintaining the physical viewing area and replacing corresponding portions thereof with virtual content, representations of physical objects within the physical viewing area remain partially or fully visible. For example, the computer system optionally replaces representations of the physical world from a topmost portion of the user's field of view downward toward a lower portion of the user's field of view, optionally partially intersecting physical objects (e.g., toys, blocks, sofas, and / or tables), such that upper portions of the physical objects are replaced by virtual content while lower portions of the representations of the physical objects continue to be displayed. In some embodiments, the representation of the physical viewing area is maintained, but visual salience is reduced (e.g., with increased translucency), at least in part because the virtual content that begins to replace the representation of the physical viewing area is displayed with reduced visual salience (e.g., with a relatively increased amount of translucency). In some embodiments, the corresponding portion of the representation of the physical viewing area (e.g., the representation of the physical object) is displayed with reduced saliency, while the other remaining portion of the physical viewing area is replaced by the virtual content. Therefore, the computer system optionally maintains the visibility of one or more portions of the user's physical environment for at least a portion of the time. In some embodiments, if one or more criteria are met, at least a portion of the representation of the physical environment is stopped from being maintained, as further described in step 830. In some embodiments, the computer system detects the presence of a physical object in the physical viewing area, and abandons replacing the representation of the physical object and / or the corresponding portion of the representation of the physical environment with virtual content. In some embodiments, if the computer system does not detect the physical object in the corresponding portion of the physical viewing area, the computer system replaces the representation of the corresponding portion of the physical viewing area with virtual content. In some embodiments, the display of the first representation of the first corresponding area including the physical object is maintained, while the representation of the second corresponding area is replaced with virtual content. The display of the representation of the physical environment while at least partially maintaining the representation of the physical environment while replacing the representation of the physical environment with virtual content visually emphasizes the presence of the physical object in the user's environment, thereby reducing potential physical collisions with such physical objects.
[0242] In some embodiments, while maintaining display of at least a portion of a representation of a corresponding area of the physical environment, such as Figure 7CAs shown, the portion of environment 702 not occupied by virtual content 716 is replaced (832) by the computer system with virtual content at an immersion level greater than the immersion threshold based on determining that the portion of the representation of the corresponding area of the physical environment has remained visible for an amount of time greater than a threshold amount of time (0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, or 15 seconds), such as by Fig.7D The area shown is replaced with virtual content 716. For example, while maintaining display of a portion of the physical viewing area, and based on determining that one or more criteria (including a criterion that is satisfied when at least the portion of the physical viewing area has remained visible for more than a threshold amount of time) are satisfied, the computer system optionally initiates replacement of the remaining portion of the physical viewing area with virtual content at an immersion level greater than an immersion threshold. In some embodiments, the replacement includes displaying an animation of the virtual content having one or more characteristics of the animation described in steps 812 and 818. Replacing at least a portion of the representation of the corresponding area of the physical environment after the representation has been visible for greater than a threshold amount of time improves the user's orientation of the physical world relative to the virtual content, thereby reducing the input of manually causing such replacement and orientation by the user, and reducing the likelihood of collision between the user and the physical environment.
[0243] In some embodiments, the representation of the corresponding area of the physical environment includes a portion of the lower area of the physical environment corresponding to the viewpoint of the user of the computer system, such as Figure 7CThe portion of the environment 702 shown that is not occupied by the virtual content 716 (834). For example, as described in step 802 and step 828. For example, during the period of maintaining the display of at least a portion of the physical viewing area described in step 828, the area corresponding to the physical viewing area of the user's viewpoint is optionally kept visible for at least a period of time. In some embodiments, the lower area corresponds to any corresponding point of the physical environment below a threshold (e.g., 0.01m, 0.025m, 0.05m, 0.25m, 0.5m, 1m, 2.5m, or 5m) height. Additionally or alternatively, the lower area optionally corresponds to an amount of the user's field of view (e.g., 0.1 degrees, 1 degrees, 3 degrees, 5 degrees, 10 degrees, 15 degrees, 30 degrees, 45 degrees, 90 degrees, or 120 degrees of the lower portion of the user's field of view). In some embodiments, the degree to which the representation of the corresponding area of the physical environment occupies the user's field of view may vary based on the orientation of the second part of the user's body (e.g., the head) relative to the physical environment. For example, when a second portion of the user's body is directed toward a boundary of the physical viewing area (e.g., the floor), the computer system optionally displays the representation of the corresponding area of the physical environment in full (e.g., without displaying the immersive virtual content or displaying a minimal amount of the immersive virtual content). In response to optionally detecting that the second of the portions has moved to a second orientation (e.g., corresponding to a field of view that includes a portion of the corresponding area of the physical environment that has been replaced by the immersive visual content), the computer system optionally simultaneously displays at least a portion of the immersive virtual content and / or the representation of the physical environment in accordance with the boundary of the virtual content displayed at an immersion level greater than an immersion threshold and the remaining portion of the physical viewing area that has not been occupied by the virtual content. Maintaining display of a lower region of the representation of the computer system user's physical environment in which the user may be moving, sitting, and / or standing while displaying the virtual content at an immersion level greater than the immersion threshold improves the user's perception of their physical surroundings, thereby reducing the likelihood of physical collisions with the environment and reducing the need to stop displaying virtual content in the lower region to obtain such perception.
[0244] In some embodiments, in response to detecting the first input via one or more input devices, the computer system displays (836) via the display generation component a selectable option that is selectable to forgo display of the visual indication in response to a future input corresponding to a request to display the virtual content at an immersion level greater than the immersion threshold, such as Fig.7DSelectable option 712-2 shown. For example, as described with respect to step 822, the computer system optionally displays multiple selectable options to indicate the user's intention to abandon displaying virtual content, such as future visual indications. The computer system optionally displays a selectable option in response to a first input to abandon the display of the visual indication in the future, and optionally detects an input that selects the selectable affordance. In some embodiments, in response to detecting a second input corresponding to a second request to display virtual content at an immersion level greater than an immersion threshold, and based on determining that one or more criteria (including criteria that were satisfied when the user of the computer system had previously selected the selectable option) are met, the computer system abandons the display of the visual indication (e.g., a geometric shape). In some embodiments, the display of the visual indication is not abandoned, but modified. For example, if one or more criteria are met, the visual indication is optionally displayed with a modified appearance (e.g., increased translucency, increased blur effect, and / or reduced brightness) in response to the second input. Presenting a selectable option to abandon the display of the visual indication later reduces the need for future input to stop the display of the visual indication.
[0245] In some embodiments, in response to detecting the first input via one or more input devices, the computer system displays (838) via the display generation component a second visual indication different from the visual indication, the second visual indication indicating that a process of determining one or more characteristics of the user's physical environment (including the corresponding area of the physical environment) has been initiated, such as Fig.7D The computer system optionally displays a progress indicator to convey that the computer system is evaluating the physical environment. In some embodiments, the progress indicator is a graphical icon (e.g., a gradually darkened and / or filled ring) that is modified according to the progress of the evaluation. In some embodiments, the progress indicator includes a grid overlaid on a representation of the physical environment that follows the outline of the physical environment (e.g., objects, floors, and / or walls). In some embodiments, the second visual indication is displayed simultaneously with the visual indication described in step 802. In some embodiments, the second visual indication is displayed until the evaluation of the physical environment is completed; in response to the completion of the evaluation, the display of the second visual indication is stopped, and the display of the visual indication is initiated. In some embodiments, one or more characteristics of the physical environment are such as the area of the floor of the physical environment, the presence of objects in the physical environment, the location of the walls in the physical environment, and / or the outline of the surface in the physical environment. Displaying an indication of the evaluation of the user's physical environment indicates that the computer system has optionally not yet responded to some user inputs, thereby reducing erroneous user inputs.
[0246] In some embodiments, when displaying the visual indication via the display generation component, the computer system displays (840a) via the display generation component selectable options that can be selected to modify the visual indication, such as Figure 7B and Figure 7B1 Selectable options 1714-1 are shown. For example, the selection option is optionally selectable to scale the shape of the visual indication in one or more directions.
[0247] In some embodiments, while displaying the selectable options via the display generation component, the computer system receives (840b) a second user input via one or more input devices, the second user input including a selection of the selectable option and a request to move the selectable option, such as input from hand 703A, such as Figure 7B and Figure 7B1 As shown. For example, the computer system optionally detects an air pinch gesture (e.g., the meeting and maintenance of contact of the index finger and thumb of the user) performed by a first part of the user (e.g., a hand) while the user's attention is directed to the selectable option and movement of the first part of the user. The second user input optionally corresponds to a selection and movement performed by a pointing device (e.g., a mouse, a stylus, and / or a glove), or another air gesture (e.g., a squeezing of the user's hand while the attention is directed to the selectable option and the visual indication is scaled according to the movement of the hand until a similar squeezing of the hand is detected).
[0248] In some embodiments, in response to receiving the second user input, the computer system modifies (840c) the visual indication based on the movement of the selectable option, such as with Figure 7C Compared with Figure 7B and Figure 7B1 704 is shown in the virtual content shown. For example, the computer system optionally detects the left and upward movement of the air pinch gesture when the user's attention is directed to the selectable option (e.g., having a semi-rectangular shape or another shape) covered on the upper left corner of the visual indication, and optionally expands the visual indication and optionally moves the selectable option according to the movement (e.g., away from the user and to the user's left or in another direction). In some embodiments, the visual indication is scaled along the corresponding one or more dimensions according to the movement. In some embodiments, the visual indication is scaled equally in all directions according to the movement. Presenting selectable options for modifying the visual indication allows the user of the computer system to reduce the visual conflict between the visual indication and the corresponding content (such as a representation of virtual content and / or a physical environment), and indicates the possible area of physical interaction to the computer system, so that the computer system can determine how to present the virtual content and modify the visual prominence of the representation of the virtual content and / or the user's physical environment accordingly.
[0249] In some embodiments, after detecting one or more characteristics of the user's physical environment (e.g., the size, shape, and / or location of one or more physical objects), a visual indication corresponding to a corresponding area of the physical environment with which the user of the computer system can interact is displayed, such as environment 702, including corresponding areas of the physical environment, such as Figure 7C The virtual content 704 (842) shown. For example, as described with respect to step 836. In some embodiments, the process is initiated and / or completed before receiving a first input requesting to display virtual content at an immersion level greater than an immersion threshold. For example, the computer system optionally initiates and / or completes the process in response to detecting that the user enters a (optionally new) physical environment (e.g., a room or other physical space). In some embodiments, the process is initiated and / or completed in response to the first input. In some embodiments, the process includes determining one or more characteristics of the physical environment or a portion of the physical environment. For example, the computer system optionally determines one or more characteristics of a first portion of the physical environment, the first portion of the physical environment optionally including a portion of the physical environment in front of the user's current viewpoint and optionally including a portion behind the user's current viewpoint (e.g., 0.01m, 0.05m, 0.1m, 0.5m, 1m, 5m, 10m, 15m, 25m, 50m, or 100m behind the user), but does not include the entire physical environment behind the user's current viewpoint. In some embodiments, the process is initiated as described in the previous embodiments and continues while the user is displaying virtual content at an immersion level greater than an immersion threshold. Displaying the visual indication after the computer system has initiated the process of determining the characteristics of the physical environment ensures that the computer system is aware of the physical environment and is thereby able to display the visual indication at corresponding areas of the physical environment corresponding to areas of possible interaction, thereby improving the user's perception of the physical environment.
[0250] It should be understood that the specific order in which the operations in method 800 are described is merely exemplary and is not intended to indicate that the described order is the only order in which the operations may be performed. A person of ordinary skill in the art will recognize many ways to reorder the operations described herein.
[0251] 9A to 9E An example of a computer system that reduces the visual prominence of immersive virtual content and displays areas of possible interaction according to some embodiments is shown.
[0252] Fig. 9A The reduction in visual prominence of virtual content according to an embodiment of the present disclosure is shown. Fig. 9AThe computer system 101 is shown displaying a three-dimensional environment 902 via a display generation component (e.g., the display generation component 120 of FIG. 1 ) from the viewpoint of a user 901 shown in a top view (e.g., facing the back wall of the physical environment in which the computer system 101 is located). Figure 6 As described above, the computer system 101 optionally includes a display generation component (e.g., a touch screen) and a plurality of image sensors (e.g., Figure 3 The image sensor 314 may include one or more of: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of a user or a portion of a user (e.g., one or more hands of a user) when the user interacts with the computer system 101. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display that includes a display generation component that displays a user interface or a three-dimensional environment to the user, and sensors that detect movement of the physical environment and / or the user's hands (such as movement interpreted by the computer system as gestures such as air gestures) (e.g., external sensors facing outward from the user), and / or sensors that detect the user's gaze (e.g., internal sensors facing inward toward the user's face).
[0253] like Fig. 9A As shown, the computer system 101 captures one or more images of the physical environment (e.g., operating environment 100) surrounding the computer system 101, including one or more objects in the physical environment surrounding the computer system 101. In some embodiments, the computer system 101 displays a representation of the physical environment in a three-dimensional environment 902, or portions of the physical environment are visible via the display generation component 120 of the computer system 101. For example, the three-dimensional environment 902 includes portions of the left and right walls, ceiling, and floor in the physical environment of the user 901.
[0254] exist Fig. 9A In the embodiment, three-dimensional environment 902 also includes virtual content, such as virtual content 904. Virtual content 916 optionally has one or more characteristics described with respect to virtual content 904, and optionally has reference FIG. 7A to FIG. 7D In some embodiments, the virtual content 916 corresponds to a virtual environment and has a reference to one or more characteristics of the virtual environment and / or immersive visual experience. FIG. 7A to FIG. 7D One or more characteristics of the virtual environment. In some embodiments, the virtual content 904 is not yet shown, or is displayed with a level of translucency that makes the virtual content 904 invisible.
[0255] In some embodiments, when as in reference method 800 and FIG. 7A to FIG. 7DWhen the virtual content 916 is displayed at an immersion level greater than the immersion threshold, the user of the computer system 101 optionally provides input to stop displaying the virtual content at an immersion level greater than the immersion threshold. Fig.7D When providing the immersive visual experience shown, computer system 101 optionally detects input including a modification of the user's viewpoint, such as movement of the user to a second location away from a corresponding location within the three-dimensional environment, such as user 901 moving toward and / or through a viewing area associated with virtual content 904 (e.g., corresponding to reference 904). FIG. 7A to FIG. 7D In some embodiments, the corresponding positioning optionally corresponds to the movement of the boundary of the virtual content 704 and the viewing area. FIG. 7A to FIG. 7D The corresponding position within the viewing area, such as the center of the viewing area, the border of the viewing area, and / or the corner of the viewing area. In some embodiments, the corresponding position has a world-locked position such that the corresponding position corresponds to a corresponding physical position in the physical environment. In some embodiments, the corresponding position corresponds to a border of the viewing area.
[0256] In some embodiments, in response to movement away from the corresponding location and / or based on determining that the modified viewpoint of the user does not correspond to the corresponding location, the computer system 101 initiates a reduction in the visual prominence of at least a portion of the virtual content 916. For example, when displaying an immersive visual experience (e.g., virtual content 916, optionally corresponding to a virtual scene of a campsite, ranch, and / or lake), in response to an input including movement of the user 901 to a second physical location (corresponding to 904) outside the viewing area, the computer system 101 begins to reduce the visual prominence of such an immersive visual experience. As described in further detail below, the reduction in visual prominence optionally includes any suitable manner of modifying the visual appearance and / or display of the virtual content 916. In some embodiments, the reduction includes stopping the display of a portion of the virtual content 916. In some embodiments, the reduction includes modifying the translucency of the portion of the virtual content 916. Reference method 1000 describes additional or alternative details about reducing the visual prominence of the representation of the virtual content 916.
[0257] Fig. 9AA modification of the visual salience of virtual content 916 according to an example of the present disclosure is shown. For example, the user 901 moves away from the viewing area corresponding to the virtual content 904 visible in the top view and is not displayed by the computer system 101. In some embodiments, the virtual content 904 is not displayed until the current viewpoint corresponds to a second position outside the corresponding area of the physical environment corresponding to the virtual content 904 (e.g., outside the viewing area). For example, when the user's position is within the viewing area, the computer system 101 optionally abandons the display of the virtual content 904, and based on determining that the user's viewpoint has shifted to the second position outside the viewing area, the computer system 101 optionally initiates the display of the virtual content 904 (e.g., corresponding to the virtual content 704). In some embodiments, the computer system 101 determines that the user has moved to the second position within the viewing area corresponding to the virtual content 904 and abandons the reduction of the visual salience of the virtual content. For example, the computer system 101 optionally detects that the user has moved to a position within the viewing area such that the user's feet have remained within the viewing area, and accordingly maintains the display of the virtual content 916.
[0258] In some embodiments, virtual content 904 corresponds to a world-locked position. For example, computer system 101 maintains an understanding of the shape and / or orientation of virtual content 904 relative to the user's physical environment, even if the user's viewpoint shifts in orientation and / or position within and / or outside the boundaries of virtual content 904. Thus, from the user's perspective, virtual content 904 optionally has a fixed positioning within environment 902, similar to a physical object such as a rug placed on the floor of environment 902.
[0259] In some embodiments, in response to an input initiating a reduction in visual prominence of virtual content 916, computer system 101 initiates a reduction in visual prominence of at least a portion of virtual content 916. For example, computer system 101 optionally detects an input comprising user 901 moving outside of a viewing area, and in response, computer system 101 optionally modifies and / or stops display of a portion of virtual content 916. In some embodiments, virtual content 916 includes one or more virtual objects (e.g., a virtual window comprising one or more user interfaces of a corresponding application such as a communication application, a media playback application, and / or a mapping application), one or more representations of virtual objects (e.g., a virtual column, a virtual car, and / or a virtual tree), and / or an immersive visual experience (e.g., an immersive visual scene such as an immersive beach, forest, and / or space scene). In some embodiments, computer system 101 interprets the movement as a request to reduce the visual prominence of at least a portion of virtual content 916, such as a request to begin viewing a portion of a physical environment and / or to completely stop display of the virtual content. In some embodiments, modifying the visual prominence of virtual content 916 to an immersion level less than the immersion threshold optionally includes ceasing display of the virtual content. For example, computer system 101 optionally ceases display of a first portion of virtual content 916, such as portion 915. Because portion 915 optionally includes virtual content displayed with reduced visual prominence (e.g., displayed with 100% transparency and / or no longer displayed), user 901 is optionally able to view physical object 910.
[0260] In some embodiments, the computer system optionally applies one or more visual effects to portion 915 of virtual content 916 to indicate a reduction in the visual prominence of portion 915 of virtual content 916. The one or more visual effects optionally include a blur effect, a feathering effect, a darkening, and / or an increase in transparency applied uniformly or non-uniformly to portion 915. For example, the leftmost area of portion 915 is optionally displayed as having a feathered edge and having a first transparency (e.g., 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 75%, 90%, or 95% transparency), and a second area to the right of the leftmost area is displayed as having a second relatively smaller transparency. In some embodiments, the non-uniformly applied visual effect includes a gradient of the visual effect (e.g., a gradient of increasing translucency from the leftmost area of portion 915 of virtual content 916 toward the rightmost area). The non-uniformity of the visual effect optionally imparts a sense of progression of the reduction in visual prominence.
[0261] In some embodiments, the direction of the reduction in visual prominence of virtual content 916 and / or portion 915 is determined based on the input of user 901. For example, because the input optionally includes a leftward movement of user 901, computer system 101 optionally initiates the reduction in visual prominence of virtual content 916 starting from the leftmost portion of virtual content 916 and / or the user's field of view (such as the left edge of the virtual content). As another example, if the input includes a movement of the user backward away from the viewing area (e.g., from virtual content 904), the computer system optionally initiates the reduction in visual prominence of virtual content 916 from an upper area of the user's field of view or a lower area of the user's field of view. Conversely, if the input optionally includes a forward movement of user 901 away from the viewing area, computer system 101 optionally initiates the reduction in visual prominence from a lower area or an upper area of the user's field of view (e.g., optionally in the opposite direction of the reduction in visual prominence in response to the backward movement). Thus, in some embodiments, the computer system 101 reduces the visual prominence of different corresponding portions of the virtual content 916 based on input including movement of the user to a second location that is outside of a corresponding area of the user's physical environment corresponding to an area where the computer system expects to interact with the virtual content (e.g., a viewing area). In addition, the computer system optionally reduces the visual prominence so that the user gains an improved understanding of how the direction of movement optionally affects the reduction in the visual prominence of the virtual content. In some embodiments, in response to detecting input including movement further away from the viewing area, the computer system 101 progressively reduces the visual prominence of larger portions of the virtual content 916. In some embodiments, in response to detecting movement of the user 901 such that the user 901 is completely outside the viewing area, the computer system 101 completely reduces the visual prominence of the virtual content 916.
[0262] Fig. 9B The display of virtual content at an immersion level greater than the immersion level is stopped. Fig. 9B, user 901 has moved to a location completely outside of the physical environment corresponding to a viewing area (e.g., corresponding to virtual content 904). In some embodiments, because all corresponding portions of the user have moved outside of the viewing area and / or because user 901 has remained outside of the viewing area for a period of time greater than a threshold amount of time (e.g., 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, 5 seconds, 10 seconds, 15 seconds, 25 seconds, 50 seconds, 100 seconds, or 500 seconds), computer system 101 has completely ceased displaying virtual content 916 at an immersion level greater than a threshold immersion level. Thus, physical objects 906 are visible without obstruction from corresponding virtual content corresponding to the immersive visual experience. In addition or in lieu of reducing the visual salience of the virtual content, computer system 101 optionally initiates display of virtual content 904 in response to the movement, such that the user optionally perceives movement to the viewing area corresponding to virtual content 904 to optionally initiate display of an immersive visual experience (e.g., as described in reference to FIG. 1 ). FIG. 7A to FIG. 7D As previously described, if the user's location does not correspond to a physical area corresponding to virtual content 904, computer system 101 optionally displays virtual content 904, as described with reference to method 800, so that the user optionally obtains a sense of a portion of the user environment in which computer system 101 expects to interact with corresponding virtual content (e.g., such as virtual content 916) associated with virtual content 904. In some embodiments, also in response to movement, computer system 101 displays virtual content 908A and virtual content 908B, which have reference to method 800. FIG. 7A to FIG. 7D One or more characteristics of virtual content 708A and / or 708B are shown.
[0263] In some embodiments, the virtual content 904 maintains a corresponding world-locked position in the three-dimensional environment 902. For example, before the user moves from an initial position at an initial viewpoint relative to his three-dimensional environment 902 to a viewing area corresponding to the virtual content 904, the virtual content 904 is optionally displayed at a first position. In response to an input corresponding to a request to stop displaying immersive virtual content (e.g., an imme...
Claims
1. A method, comprising: At a computer system in communication with a display generating component and one or more input devices: detecting, via the one or more input devices, a first input corresponding to a request to display virtual content that will visually replace a portion of a representation of a physical environment in which a user of the computer system is located while using the computer system; In response to detecting the first input via the one or more input devices and in accordance with determining that the first input corresponds to a request to display the virtual content at an immersion level greater than an immersion threshold: displaying, via the display generation component, a visual indication corresponding to a corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold, while a representation of the corresponding area of the physical environment is visible via the display generation component; and Displaying, via the display generation component, the virtual content at the immersion level greater than the immersion threshold, includes replacing at least a portion of the representation of the corresponding area of the physical environment with the virtual content after displaying the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold.
2. The method of claim 1 , wherein displaying, via the display generation component, the visual indication corresponding to the respective area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold, comprises: Based on determining that the user is located at a first location in the physical environment: The visual indication corresponding to the corresponding area is a first visual indication corresponding to a first area of the physical environment, and Replacing the portion of the representation of the corresponding area of the physical environment with the virtual content comprises replacing at least a portion of the representation of the first area of the physical environment with the virtual content; and Based on determining that the user is located at a second location in the physical environment that is different from the first location: the visual indication corresponding to the corresponding area is a second visual indication corresponding to a second area of the physical environment, the second area being different from the first area of the physical environment, and Replacing the portion of the representation of the corresponding area of the physical environment with the virtual content includes replacing at least a portion of the representation of the second area of the physical environment with the virtual content.
3. The method of any one of claims 1 to 2, wherein the visual indication corresponding to the respective area of the physical environment with which a user of the computer system is likely to interact is displayed in association with a floor of the physical environment. The method of claim 3 , wherein the visual indication has a first shape, and wherein the visual indication is at least partially translucent.
5. The method of any one of claims 3 to 4, wherein the first shape is elliptical and has a first corresponding diameter, and the visual indication comprises a plurality of shapes, the plurality of shapes comprising the first shape and a second shape different from the first shape, wherein the second shape has a second diameter different from the first diameter.
6. A method according to any one of claims 1 to 5, wherein the displaying of the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system may interact via the display generating component includes displaying an animation of a boundary of the visual indication extending from a first position in the three-dimensional environment corresponding to the corresponding part of the user to a second position in the three-dimensional environment different from the first position.
7. The method of claim 6, wherein the animation comprises a visual effect applied to a surface of the corresponding area of the physical environment with which the user of the computer system is likely to interact.
8. The method according to any one of claims 6 to 7, further comprising: When the boundary of the visual indication corresponding to the respective area of the physical environment is animated via the display generation component, the animation of the boundary of the visual indication is stopped based on determining that the visual indication has been animated for a time period greater than a threshold time period.
9. The method of any one of claims 1 to 8, wherein displaying the virtual content at the immersion level greater than the immersion threshold via the display generation component comprises: replacing a second representation of a second corresponding area of the representation of the physical environment with the first portion of the virtual content, the second corresponding area corresponding to an upper area of the user's viewpoint; as well as After replacing the second representation of the second corresponding area, replacing a third representation of a third corresponding area of the representation of the physical environment with a second portion of the virtual content, the third corresponding area corresponding to a lower area of the user's viewpoint, the lower area of the user's viewpoint being lower than the upper area of the user's viewpoint.
10. The method according to any one of claims 1 to 9, further comprising: In response to detecting the first input via the one or more input devices, and based on determining that the one or more criteria are satisfied, displaying, via the display generation component, corresponding virtual content indicating that the virtual content is to be displayed at the immersion level greater than the immersion threshold, wherein the one or more criteria include a criterion that is satisfied when the first input corresponds to the request to display the virtual content at the immersion level greater than the immersion threshold. 11 . The method of claim 10 , wherein satisfying the one or more criteria is independent of a number of times virtual content has been displayed at the immersion level greater than the immersion threshold.
12. The method according to any one of claims 10 to 11, further comprising: In response to detecting the first input via the one or more input devices, and based on determining that the one or more criteria are not satisfied, displaying the corresponding virtual content via the display generation component is foregone.
13. The method of claim 12, wherein the one or more criteria include a criterion that is satisfied based on a recency of a previous interaction by the user of the computer system with the virtual content.
14. A method according to any one of claims 12 to 13, wherein the one or more criteria include a criterion that is satisfied based on the recency of a corresponding input previously received detected via the one or more input devices, the corresponding input corresponding to a request to display corresponding virtual content in the corresponding area of the physical environment at an immersion level greater than the immersion threshold.
15. A method according to any one of claims 1 to 14, wherein after displaying the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system may interact, replacing at least a portion of the representation of the corresponding area of the physical environment with the virtual content includes maintaining display of at least a portion of the representation of the corresponding area of the physical environment.
16. The method according to claim 15, further comprising: While maintaining the display of at least a portion of the representation of the corresponding area of the physical environment, based on determining that the portion of the representation of the corresponding area of the physical environment has remained visible for an amount of time greater than a threshold amount of time, replacing at least a portion of the representation of the corresponding area of the physical environment with the virtual content having the immersion level greater than the immersion threshold.
17. The method of any one of claims 15 to 16, wherein the representation of the corresponding area of the physical environment comprises a portion of the physical environment corresponding to a lower area of a viewpoint of the user of the computer system.
18. The method according to any one of claims 1 to 17, further comprising: In response to detecting the first input via the one or more input devices, displaying a selectable option via the display generation component, the selectable option being selectable to forgo the display of the visual indication in response to a future input corresponding to a request to display the virtual content at the immersion level greater than the immersion threshold.
19. The method according to any one of claims 1 to 18, further comprising: In response to detecting the first input via the one or more input devices, a second visual indication different from the visual indication is displayed via the display generating component, and the second visual indication indicates that a process for determining one or more characteristics of the physical environment of the user, including the corresponding area of the physical environment, has been initiated.
20. The method according to any one of claims 1 to 19, further comprising: while displaying the visual indication via the display generating component, displaying via the display generating component a selectable option that can be selected to modify the visual indication; receiving, via the one or more input devices, a second user input while displaying the selectable option via the display generating component, the second user input comprising a selection of the selectable option and a request to move the selectable option; as well as In response to receiving the second user input, the visual indication is modified according to the movement of the selectable option.
21. A method according to any one of claims 1 to 20, wherein after detecting one or more characteristics of the user's physical environment, the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system is likely to interact is displayed, and the physical environment includes the corresponding area of the physical environment.
22. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for: detecting, via the one or more input devices, a first input corresponding to a request to display virtual content that will visually replace a portion of a representation of a physical environment in which a user of the computer system is located while using the computer system; In response to detecting the first input via the one or more input devices and in accordance with determining that the first input corresponds to a request to display the virtual content at an immersion level greater than an immersion threshold: displaying, via the display generation component, a visual indication corresponding to a corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold, while a representation of the corresponding area of the physical environment is visible via the display generation component; and Displaying, via the display generation component, the virtual content at the immersion level greater than the immersion threshold, includes replacing at least a portion of the representation of the corresponding area of the physical environment with the virtual content after displaying the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold.
23. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising: detecting, via the one or more input devices, a first input corresponding to a request to display virtual content that will visually replace a portion of a representation of a physical environment in which a user of the computer system is located while using the computer system; In response to detecting the first input via the one or more input devices and in accordance with determining that the first input corresponds to a request to display the virtual content at an immersion level greater than an immersion threshold: displaying, via the display generation component, a visual indication corresponding to a corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold, while a representation of the corresponding area of the physical environment is visible via the display generation component; and Displaying, via the display generation component, the virtual content at the immersion level greater than the immersion threshold, includes replacing at least a portion of the representation of the corresponding area of the physical environment with the virtual content after displaying the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold.
24. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; means for detecting, via the one or more input devices, a first input corresponding to a request to display virtual content that will visually replace a portion of a representation of a physical environment in which a user of the computer system is located while using the computer system; Means for, in response to detecting the first input via the one or more input devices and in accordance with determining that the first input corresponds to a request to display the virtual content at an immersion level greater than an immersion threshold: displaying, via the display generation component, a visual indication corresponding to a corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold, while a representation of the corresponding area of the physical environment is visible via the display generation component; and Displaying, via the display generation component, the virtual content at the immersion level greater than the immersion threshold, includes replacing at least a portion of the representation of the corresponding area of the physical environment with the virtual content after displaying the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system is likely to interact when the virtual content is displayed at the immersion level greater than the immersion threshold.
25. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for executing any one of the methods according to claims 1 to 21.
26. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform any of the methods of claims 1 to 21.
27. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and Apparatus for carrying out any of the methods according to claims 1 to 21.
28. A method comprising: At a computer system in communication with one or more input devices and a display generating component: displaying, via the display generation component, virtual content having a world-locked position relative to a physical environment visible when the virtual content is displayed, wherein the virtual content is displayed from a first viewpoint of a user of the computer system, wherein the first viewpoint corresponds to a first physical location within a corresponding area of the physical environment associated with viewing the virtual content; detecting, via the one or more input devices, movement of the user to a second physical location in the physical environment that is different from the first physical location in the physical environment while displaying the virtual content via the display generation component; and In response to detecting the movement of the user to the second physical location in the physical environment and based on determining that the second physical location is outside of the corresponding area in the physical environment associated with viewing the virtual content: reducing the visual prominence of at least a portion of the virtual content; as well as A visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed via the display generation component, wherein the visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed at a location in the physical environment corresponding to the corresponding area.
29. The method according to claim 28, further comprising: In response to detecting the movement of the user to the second physical location in the physical environment and based on determining that the second physical location is at least partially within the corresponding area in the physical environment associated with viewing the virtual content: Displaying, via the display generation component, the visual indication of the corresponding area in the physical environment associated with viewing the virtual content is foregone.
30. The method according to any one of claims 28 to 29, further comprising: In response to detecting the movement of the user to the second physical location in the physical environment and based on determining that the second physical location is at least partially within the corresponding area in the physical environment associated with viewing the virtual content: Reducing the visual prominence of the at least a portion of the virtual content is abandoned.
31. A method according to any one of claims 28 to 30, wherein the position in the physical environment corresponding to the respective area where the visual indication is displayed is a first respective world-locked position relative to the physical environment, the method further comprising: while displaying, via the display generating component, the visual indication at the first corresponding world-locked position relative to the physical environment, detecting, via the one or more input devices, an input corresponding to a request to move the visual indication to a second corresponding world-locked position relative to the physical environment, wherein the second corresponding world-locked position is different from the first corresponding world-locked position, and In response to detecting, via the one or more input devices, the input corresponding to the request to move the visual indication of the corresponding area in the physical environment associated with viewing the virtual content: The visual indication of the corresponding area in the physical environment associated with viewing the virtual content at the second corresponding world-locked position is displayed via the display generation component.
32. The method according to any one of claims 28 to 31, further comprising: while displaying, via the display generating component, the visual indication at the location in the physical environment corresponding to the corresponding area and when the user is at the second physical location in the physical environment, wherein the second physical location is outside of the corresponding area in the physical environment associated with viewing the virtual content, detecting, via the one or more input devices, a second movement of the user to a third physical location in the physical environment different from the second physical location; as well as In response to detecting the second movement of the user to the third physical location in the physical environment, the visual prominence of at least a portion of the virtual content is increased based on determining that the third physical location is at least partially within the corresponding area in the physical environment associated with viewing the virtual content.
33. The method of any one of claims 28 to 32, wherein the visual indication of the corresponding area in the physical environment associated with viewing the virtual content includes information associated with the virtual content.
34. The method of claim 33, wherein the information includes an indication of an application associated with the virtual content.
35. The method of claim 34, wherein the indication of the application comprises a visual representation of the application displayed at a corresponding location corresponding to the location in the physical environment and above a floor of the physical environment, wherein The location in the physical environment corresponds to the corresponding area in the physical environment associated with viewing the virtual content.
36. A method according to any one of claims 33 to 35, wherein the information includes a visual representation having a corresponding shape displayed at a corresponding area of the physical environment, the corresponding area corresponding to the location in the physical environment corresponding to the corresponding area.
37. The method of claim 36, wherein the corresponding shape is displayed with a visual characteristic having a corresponding value based on an application associated with the virtual content.
38. The method of claim 37, wherein the corresponding value is based on corresponding content associated with the application.
39. A method according to any one of claims 37 to 38, wherein the corresponding value is based on a color included in the visual representation of the application.
40. The method of claim 34, wherein the information comprises a simulated lighting effect displayed at a corresponding area of the physical environment, the corresponding area of the physical environment corresponding to the location in the physical environment corresponding to the corresponding area.
41. The method according to any one of claims 28 to 40, further comprising: detecting, when the user is at the second physical location, wherein the second physical location is outside of the corresponding area in the physical environment associated with viewing the virtual content, and when the visual indication is displayed, via the display generation component, at the location in the physical environment corresponding to the corresponding area, an input corresponding to a request to re-center the corresponding virtual content based on a current viewpoint of the user, via the one or more input devices; as well as In response to detecting, via the one or more input devices, the input corresponding to the request to re-center corresponding virtual content based on the current viewpoint of the user, the visual prominence of the at least a portion of the virtual content is increased.
42. The method of claim 41, further comprising: In response to detecting, via the one or more input devices, the input corresponding to the request to re-center the corresponding virtual content based on the current viewpoint of the user, displaying, via the display generation component, a second visual indication different from the visual indication corresponding to the corresponding area of the physical environment with which the user of the computer system is likely to interact, wherein the display includes displaying an animation of the second visual indication appearing at a location corresponding to the second physical location of the user.
43. The method of any one of claims 28 to 42, wherein reducing the visual prominence of the at least a portion of the virtual content comprises: initiating the reducing the visual prominence of the at least a portion of the virtual content from a first direction based on determining that the movement of the user to the second physical location is in a first direction relative to the physical environment; as well as Based on determining that the movement of the user to the second physical location is in a second direction relative to the physical environment that is different from the first direction, the reducing the visual prominence of the at least a portion of the virtual content is initiated from the second direction.
44. The method of any one of claims 28 to 43, wherein the displaying, via the display generation component, the visual indication of the corresponding area in the physical environment associated with viewing the virtual content comprises displaying the visual indication at a first size relative to the physical environment, the method further comprising: While displaying the visual indication of the corresponding area in the physical environment associated with viewing the virtual content, and while the location of the user of the computer system is the second location in the physical environment outside of the corresponding area in the physical environment: Based on determining that one or more criteria are satisfied, the one or more criteria including a criterion satisfied when the position of the user of the computer system has remained outside of the corresponding area in the physical environment for a threshold amount of time, changing the size of the visual indication from the first size to a second size relative to the physical environment, wherein the second size is smaller than the first size.
45. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for: displaying, via the display generation component, virtual content having a world-locked position relative to a physical environment visible when the virtual content is displayed, wherein the virtual content is displayed from a first viewpoint of a user of the computer system, wherein the first viewpoint corresponds to a first physical position within a corresponding area in the physical environment associated with viewing the virtual content; detecting, via the one or more input devices, movement of the user to a second physical location in the physical environment that is different from the first physical location in the physical environment while displaying the virtual content via the display generation component; and In response to detecting the movement of the user to the second physical location in the physical environment and based on determining that the second physical location is outside of the corresponding area in the physical environment associated with viewing the virtual content: reducing the visual prominence of at least a portion of the virtual content; as well as A visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed via the display generation component, wherein the visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed at a location in the physical environment corresponding to the corresponding area.
46. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising: displaying, via the display generation component, virtual content having a world-locked position relative to a physical environment visible when the virtual content is displayed, wherein the virtual content is displayed from a first viewpoint of a user of the computer system, wherein the first viewpoint corresponds to a first physical position within a corresponding area in the physical environment associated with viewing the virtual content; detecting, via the one or more input devices, movement of the user to a second physical location in the physical environment that is different from the first physical location in the physical environment while displaying the virtual content via the display generation component; and In response to detecting the movement of the user to the second physical location in the physical environment and based on determining that the second physical location is outside of the corresponding area in the physical environment associated with viewing the virtual content: reducing the visual prominence of at least a portion of the virtual content; and A visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed via the display generation component, wherein the visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed at a location in the physical environment corresponding to the corresponding area.
47. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; means for displaying, via the display generation component, virtual content having a world-locked position relative to a physical environment visible when the virtual content is displayed, wherein the virtual content is displayed from a first viewpoint of a user of the computer system, wherein the first viewpoint corresponds to a first physical position within a respective area of the physical environment associated with viewing the virtual content; means for detecting, via the one or more input devices, movement of the user to a second physical location in the physical environment while displaying the virtual content via the display generation component, the second physical location being different from the first physical location in the physical environment; and Means for, in response to detecting the movement of the user to the second physical location in the physical environment and in accordance with determining that the second physical location is outside of the corresponding area in the physical environment associated with viewing the virtual content: reducing the visual prominence of at least a portion of the virtual content; and A visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed via the display generation component, wherein the visual indication of the corresponding area in the physical environment associated with viewing the virtual content is displayed at a location in the physical environment corresponding to the corresponding area.
48. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing any of the methods according to claims 28 to 44.
49. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform any of the methods of claims 28 to 44.
50. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and Apparatus for carrying out any of the methods according to claims 28 to 44.
51. A method comprising: At a computer system in communication with one or more input devices and one or more output generating components including a display generating component: while displaying first virtual content via the display generation component, wherein the first virtual content is obscuring a first portion of a physical environment of a user of the computer system, generating via the one or more output generation components a warning based on determining that a first physical object located at a first position in the first portion of the physical environment conflicts with a potential range of motion of the user in the physical environment, wherein the alert indicates the first position of the first physical object; as well as detecting behavior of the user when generating the alert via the one or more output generating components; and In response to detecting the behavior of the user: reducing the prominence of the alert based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated meets one or more criteria; as well as Based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated does not meet the one or more criteria, reducing the significance of the alert is forgone.
52. The method of claim 51, further comprising: In response to detecting the behavior of the user: The prominence of the alert is increased based on a determination that the detected behavior of the user of the computer system that was detected when the alert was generated does not satisfy the one or more criteria.
53. The method of any one of claims 51 to 52, wherein the one or more criteria include a criterion that is satisfied when a gaze of the user of the computer system is directed toward the alert.
54. A method according to any one of claims 51 to 53, wherein the one or more criteria include a criterion that is satisfied when the behavior of the user reduces the conflict of the first physical object with the potential range of motion of the user in the physical environment.
55. The method of claim 54, wherein the behavior reduces the conflict of the first physical object with the potential range of motion of the user in the physical environment when the user's movement speed toward the first physical object decreases.
56. The method of any one of claims 54 to 55, wherein the behavior reduces the conflict of the first physical object with the potential range of motion of the user in the physical environment when the user's movement toward the first physical object stops.
57. The method of any one of claims 54 to 56, wherein the behavior reduces the conflict of the first physical object with the potential range of motion of the user in the physical environment as the user moves away from the first physical object.
58. The method of any one of claims 51 to 57, wherein generating the warning in accordance with the determining that the first physical object located at the first position in the first portion of the physical environment conflicts with the potential range of motion of the user in the physical environment comprises generating the warning having a first salience, the method further comprising: In response to detecting the behavior of the user: Increasing the significance of the alert from the first significance to a second significance greater than the first significance based on determining that a detected behavior of the user of the computer system that was detected when the alert was generated does not meet the one or more criteria due to the detected behavior increasing the conflict of the first physical object with the potential range of motion of the user in the physical environment.
59. A method according to any one of claims 51 to 58, wherein the one or more criteria include a criterion that is satisfied when the detected behavior reduces the conflict of the first physical object with the potential range of motion of the user in the physical environment.
60. The method of any one of claims 51 to 59, wherein generating the alert comprises displaying, via the display generation component, second virtual content separate from the first virtual content, wherein the second virtual content is not displayed prior to generating the alert.
61. A method according to any one of claims 51 to 60, wherein generating the warning includes reducing the visual prominence of a first portion of the first virtual content corresponding to the first position of the first physical object relative to a second portion of the first virtual object corresponding to a second position in the physical environment, so that the first physical object is at least partially visible through the first portion of the first virtual content.
62. The method of any one of claims 51 to 61, wherein a first audio output associated with the first virtual content is generated while displaying the first virtual content, and generating the alert comprises changing one or more characteristics of the first audio output.
63. The method of claim 62, wherein changing the one or more characteristics of the first audio output comprises generating a second audio output having a directional characteristic corresponding to the first position of the first physical object.
64. A method according to any one of claims 62 to 63, wherein changing the one or more characteristics of the first audio output includes reducing the auditory significance of a corresponding portion of the first audio output, wherein the corresponding portion of the first audio output has a directional characteristic corresponding to the first position of the first physical object.
65. The method of any one of claims 51 to 64, wherein generating the alert comprises: A virtual lighting effect is displayed via the display generation component from a direction of the display generation component corresponding to the first position of the first physical object.
66. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for: while displaying first virtual content via the display generation component, wherein the first virtual content is obscuring a first portion of a physical environment of a user of the computer system, generating via the one or more output generation components a warning based on determining that a first physical object located at a first position in the first portion of the physical environment conflicts with a potential range of motion of the user in the physical environment, wherein the warning indicates the first position of the first physical object; as well as detecting behavior of the user when generating the alert via the one or more output generating components; and In response to detecting the behavior of the user: reducing the prominence of the alert based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated meets one or more criteria; as well as Based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated does not meet the one or more criteria, reducing the significance of the alert is forgone.
67. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising: while displaying first virtual content via the display generation component, wherein the first virtual content is obscuring a first portion of a physical environment of a user of the computer system, generating via the one or more output generation components a warning based on determining that a first physical object located at a first position in the first portion of the physical environment conflicts with a potential range of motion of the user in the physical environment, wherein the warning indicates the first position of the first physical object; as well as detecting behavior of the user when generating the alert via the one or more output generating components; and In response to detecting the behavior of the user: reducing the prominence of the alert based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated meets one or more criteria; as well as Based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated does not meet the one or more criteria, reducing the significance of the alert is forgone.
68. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; means for generating, via the one or more output generating components, a warning based on determining that a first physical object located at a first position in the first portion of the physical environment conflicts with a potential range of motion of the user in the physical environment while displaying, via the display generating component, first virtual content, wherein the first virtual content is obscuring a first portion of a physical environment of a user of the computer system, wherein the warning indicates the first position of the first physical object; and means for detecting behavior of said user when generating said alert via said one or more output generating components; and Means for performing the following operations in response to detecting the behavior of the user: reducing the prominence of the alert based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated meets one or more criteria; as well as Based on determining that the detected behavior of the user of the computer system that was detected when the alert was generated does not meet the one or more criteria, reducing the significance of the alert is forgone.
69. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods according to claims 51 to 65.
70. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform any of the methods described in claims 51 to 65.
71. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and Apparatus for performing any of the methods according to claims 51 to 65.
72. A method comprising: At a computer system in communication with a display generating component and one or more input devices: While displaying first virtual content via the display generation component, wherein the first virtual content is obscuring a first portion of a physical environment, detecting via the one or more input devices a first person located in the first portion of the physical environment; as well as In response to detecting the first person in the first portion of the physical environment: increasing the visual prominence of the first person relative to the first virtual content based on determining that the first person satisfies one or more criteria, wherein the one or more criteria indicate that the computer system has detected that the attention of the first person is directed toward a user of the computer system; as well as Based on determining that the first person does not meet the one or more criteria, increasing the visual prominence of the first person relative to the first virtual content is abandoned.
73. The method of claim 72, wherein increasing the visual prominence of the first person relative to the first virtual content comprises increasing the visual prominence of the first person to a first visual prominence relative to the first virtual content, the method further comprising: In response to detecting the first person in the first portion of the physical environment and before the first person satisfies the one or more criteria, increasing the visual prominence of the first person relative to the first virtual content to a second visual prominence relative to the first virtual content, wherein the second visual prominence is less than the first visual prominence.
74. The method of any one of claims 72 to 73, wherein the one or more criteria include a criterion that is satisfied when the computer system has detected that the gaze of the first person is directed toward the user of the computer system.
75. A method according to any one of claims 72 to 74, wherein the one or more criteria include a criterion that is satisfied when the computer system has detected the first person's speech that satisfies one or more second criteria.
76. A method according to any one of claims 72 to 75, wherein the one or more criteria include a criterion that is satisfied when the computer system has detected a corresponding part of the first person's body that satisfies one or more second criteria.
77. A method according to claim 76, wherein the criterion is met when the computer system detects that the distance between the corresponding part of the body of the first person and the user of the computer system is less than a threshold distance.
78. A method according to any one of claims 76 to 77, wherein the criterion is met when the computer system detects that the orientation of the corresponding part of the body of the first person relative to the user of the computer system is within a threshold orientation.
79. The method of any one of claims 72 to 78, further comprising: while displaying the first virtual content via the display generation component, detecting, via the one or more input devices, a respective person located in a respective portion of the physical environment that is being obscured by the first virtual content; as well as In response to detecting the respective person in the respective portion of the physical environment, and in accordance with determining that the respective person satisfies the one or more criteria: In response to determining that the first setting of the computer system has a first value, increasing the visual prominence of the corresponding person relative to the first virtual content; and Based on determining that the first setting of the computer system has a second value different than the first value, increasing the visual prominence of the corresponding person relative to the first virtual content is abandoned.
80. The method of claim 79, further comprising: A control user interface of the computer system is displayed via the display generation component, the control user interface including a selectable option selectable to set the first value or the second value for the first setting.
81. A method according to any one of claims 72 to 80, wherein increasing the visual prominence of the first person relative to the first virtual content includes modifying the visual appearance of a corresponding portion of the first virtual content, wherein a shape of the corresponding portion of the first virtual content is asymmetric along at least one axis.
82. The method of any one of claims 72 to 81, wherein increasing the visual prominence of the first person relative to the first virtual content comprises modifying a visual appearance of a corresponding portion of the first virtual content, the method further comprising: When the first person satisfies the one or more first criteria and when the first person has an increased visual prominence relative to the first virtual content: detecting, via the one or more input devices, movement of the first person from the first position relative to the first virtual content to a second position relative to the first virtual content that is different from the first position, when the corresponding portion of the first virtual content is a first corresponding portion of the first virtual content corresponding to a first position of the first person relative to the first virtual content; as well as In response to detecting the movement of the first person from the first position relative to the first virtual content to the second position relative to the first virtual content, modifying the visual appearance of a second corresponding portion of the first virtual content corresponding to the second position of the first person relative to the first virtual content.
83. The method of any one of claims 72 to 82, wherein increasing the visual prominence of the first person relative to the first virtual content comprises: increasing the visual prominence of the first person relative to the first virtual content to a first visual prominence relative to the first virtual content; as well as After the visual prominence of the first person relative to the first virtual content is increased to the first visual prominence, the visual prominence of the first person relative to the first virtual content is gradually reduced from the first visual prominence to a second visual prominence relative to the first virtual content.
84. The method of claim 83, further comprising: detecting, via the one or more input devices, attention of the user of the computer system directed toward the first person when the visual salience of the first person relative to the first virtual content is the second visual salience relative to the first virtual content; as well as In response to detecting the attention of the user of the computer system directed toward the first person, the visual prominence of the first person relative to the first virtual content is increased to a third visual prominence relative to the first virtual content, wherein the third visual prominence is greater than the second visual prominence.
85. A method according to any one of claims 83 to 84, wherein the one or more criteria are satisfied based on a level of attention detected by the computer system being greater than a threshold level of attention.
86. The method of any one of claims 72 to 85, wherein increasing the visual prominence of the first person relative to the first virtual content comprises increasing the visual prominence of the first person relative to the first virtual content to a first visual prominence relative to the first virtual content, the method further comprising: while displaying the first virtual content via the display generation component, detecting, via the one or more input devices, a corresponding person located in a corresponding portion of the physical environment obscured by the first virtual content; as well as In response to detecting the corresponding person in the corresponding portion of the physical environment, and based on determining that the corresponding person does not meet the one or more criteria, the visual prominence of the corresponding person relative to the first virtual content is increased to a second visual prominence relative to the first virtual content, and the second visual prominence is less than the first visual prominence relative to the first virtual content.
87. The method of any one of claims 72 to 86, wherein increasing the visual prominence of the first person relative to the first virtual content comprises increasing the visual prominence of the first person relative to the first virtual content to a first visual prominence relative to the first virtual content, the method further comprising: detecting input from the user of the computer system via the one or more input devices when the first person has the first visual saliency relative to the first virtual content; as well as Upon detecting the input from the user of the computer system, and based on determining that the input from the user of the computer system satisfies one or more second criteria, reducing the visual prominence of the first person relative to the first virtual content.
88. The method of claim 87, wherein the one or more second criteria include a criterion that is satisfied when the input from the user includes an input for moving the first virtual content.
89. The method of any one of claims 87 to 88, wherein the one or more second criteria include a criterion that is satisfied when the input from the user includes an input to scroll through the first virtual content.
90. The method of any one of claims 87 to 89, wherein the one or more second criteria include criteria that are satisfied when the input from the user includes input interacting with one or more controls associated with the first virtual content.
91. A method according to any one of claims 87 to 90, wherein the one or more second criteria include a criterion that is satisfied when the input from the user includes a part of the user's body in a corresponding posture.
92. The method of any one of claims 72 to 91, wherein the first virtual content is visible via the display generation component simultaneously with a corresponding portion of an environment, the method further comprising: detecting, via the one or more input devices, attention of the user of the computer system directed toward the first person while the corresponding portion of the environment is visible at a first visual salience relative to the environment; as well as In response to detecting the user's attention directed toward the first person, a visual prominence of the corresponding portion of the environment is increased to a second visual prominence relative to the environment.
93. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for: While displaying first virtual content via the display generation component, wherein the first virtual content is obscuring a first portion of a physical environment, detecting via the one or more input devices a first person located in the first portion of the physical environment; as well as In response to detecting the first person in the first portion of the physical environment: increasing the visual prominence of the first person relative to the first virtual content based on determining that the first person satisfies one or more criteria, wherein the one or more criteria indicate that the computer system has detected that the attention of the first person is directed toward a user of the computer system; as well as Based on determining that the first person does not meet the one or more criteria, increasing the visual prominence of the first person relative to the first virtual content is abandoned.
94. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising: While displaying first virtual content via the display generation component, wherein the first virtual content is obscuring a first portion of a physical environment, detecting via the one or more input devices a first person located in the first portion of the physical environment; as well as In response to detecting the first person in the first portion of the physical environment: increasing the visual prominence of the first person relative to the first virtual content based on determining that the first person satisfies one or more criteria, wherein the one or more criteria indicate that the computer system has detected that the attention of the first person is directed toward a user of the computer system; as well as Based on determining that the first person does not meet the one or more criteria, increasing the visual prominence of the first person relative to the first virtual content is abandoned.
95. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; means for detecting, via the one or more input devices, a first person located in the first portion of the physical environment while displaying, via the display generation component, first virtual content, wherein the first virtual content is obscuring the first portion of the physical environment; and Means for, in response to detecting the first person in the first portion of the physical environment: increasing the visual prominence of the first person relative to the first virtual content based on determining that the first person satisfies one or more criteria, wherein the one or more criteria indicate that the computer system has detected that the attention of the first person is directed toward a user of the computer system; as well as Based on determining that the first person does not meet the one or more criteria, increasing the visual prominence of the first person relative to the first virtual content is abandoned.
96. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for executing any of the methods according to claims 72 to 92.
97. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, cause the computer system to perform any of the methods described in claims 72 to 92.
98. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; Memory; and Apparatus for performing any of the methods of claims 72 to 92.