Masked objects in three-dimensional environment
By reducing the visual prominence of virtual content and improving interactive feedback in the three-dimensional environment, the problems of low interaction efficiency and complex operation in the prior art are solved, and a more intuitive and efficient user interaction experience is achieved.
Patent Information
- Application Number
- CN202510275473.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-20
- Filing Date
- 2023-04-20
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the methods and interfaces used to interact with content in a three-dimensional environment have problems such as low efficiency, insufficient feedback, complex operation and error-prone, resulting in a large cognitive burden on users and wasted energy consumption.
By providing improved user interfaces and interaction methods, the number and complexity of user input is reduced and the efficiency of the human-computer interface is improved. Specific measures include reducing the visual prominence of virtual content in a three-dimensional environment, enabling physical objects to penetrate virtual content, and improving the user's interactive experience through visual and audio feedback.
It achieves more efficient and intuitive user interaction, reduces operating time and energy consumption, and improves the experience of virtual/augmented reality environments.
Smart Images

Figure CN120215703A_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese patent application with an application date of April 20, 2023, an application number of 202380048411.5, and an invention title of "Occluded Objects in a Three-Dimensional Environment". Technical Field
[0002] The present invention generally relates to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Art
[0003] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention
[0004] Some methods and interfaces for interacting with an environment that includes at least some virtual elements (e.g., an application, an augmented reality environment, a mixed reality environment, and a virtual reality environment) are cumbersome, inefficient, and limited. For example, systems that provide inadequate feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where virtual object manipulation is complex, cumbersome, and error-prone impose a significant cognitive burden on users and detract from the experience of the virtual / augmented reality environment. In addition, these methods take longer than necessary, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0005] Accordingly, there is a need for computer systems with improved methods and interfaces to provide computer-generated experiences to users such that the interaction of the users with the computer systems is more effective and intuitive for the users. Such methods and interfaces optionally supplement or replace conventional methods for providing an extended reality experience. Such methods and interfaces form a more effective human-machine interface by helping users understand the connection between the inputs provided and the device's response to these inputs, thereby reducing the quantity, degree, and / or nature of the inputs from the users.
[0006] The above-described deficiencies and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to a display generation component, the computer system also has one or more output devices, which include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or sets of instructions stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or the computer system) or the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gaming, making phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured to be executed by one or more processors.
[0007] There is a need for electronic devices having improved methods and interfaces for interacting with content in a three-dimensional environment. Such methods and interfaces can supplement or replace conventional methods for interacting with content in a three-dimensional environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.
[0008] In some embodiments, according to some embodiments, a computer system reduces the visual prominence of virtual content relative to a three-dimensional environment. In some embodiments, according to some embodiments, the computer system displays an indication of a physical object in the three-dimensional environment.
[0009] Note that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described in this specification are not comprehensive. Specifically, many additional features and advantages will be apparent to those of ordinary skill in the art from the drawings, the specification, and the claims. In addition, it should be noted that the language used in this specification has been selected for readability and guidance purposes and may not have been selected to depict or define the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals indicate corresponding parts in all the drawings.
[0011] Figure 1 is a block diagram illustrating an operating environment of a computer system for providing an XR experience according to some embodiments.
[0012] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate an XR experience for a user according to some embodiments.
[0013] Figure 3 is a block diagram illustrating a display generation component of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.
[0014] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input according to some embodiments.
[0015] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.
[0016] Figure 6 is a flowchart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.
[0017] Figures 7A to 7E illustrates an example of a computer system that reduces the visual prominence of virtual content relative to a three-dimensional environment according to some embodiments.
[0018] Figures 8A to 8IFIG. is a flow diagram illustrating an exemplary method for reducing the visual prominence of virtual content relative to a three-dimensional environment, according to some embodiments.
[0019] Figures 9A to 9E An example of a computer system for displaying an indication of a physical object in a three-dimensional environment, according to some embodiments, is illustrated.
[0020] Figures 10A to 10F FIG. is a flow diagram illustrating a method for displaying an indication of a physical object in a three-dimensional environment, according to some embodiments. DETAILED DESCRIPTION
[0021] According to some embodiments, the present disclosure relates to a user interface for providing a computer-generated reality (CGR) experience to a user.
[0022] The systems, methods, and GUIs described herein provide improved ways for an electronic device to facilitate interaction with and manipulation of objects in a three-dimensional environment.
[0023] In some embodiments, a three-dimensional environment including first virtual content is visible via a display generation component and obscures one or more portions of the physical environment. In some embodiments, a first physical object is detected within the physical environment, and based on determining that the first physical object meets one or more criteria to constitute a significant social interaction, the computer system reduces the visual prominence of the first virtual content and enables the first physical object to penetrate the first virtual content. However, if a user-defined setting is set to prevent penetration of physical objects that meet one or more criteria, the computer system optionally maintains the visual prominence of the first virtual content.
[0024] In some embodiments, a three-dimensional environment including first virtual content is visible via a display generation component and obscures one or more portions of the physical environment. In some embodiments, a first physical object is detected within the physical environment, and based on determining that the first physical object meets one or more criteria to constitute a significant social interaction, the computer system reduces the visual prominence of the first virtual content and enables the first physical object to penetrate the first virtual content. However, if the first physical object does not meet one or more criteria, an indication representing the first physical object is optionally displayed concurrently with the first virtual content.
[0025] Figures 1 to 6 A description of an example computer system for providing an XR experience to a user (such as those described below with respect to methods 800 and 1000) is provided. Figures 7A to 7E An example of a computer system for reducing the visual prominence of virtual content relative to a three-dimensional environment, according to some embodiments, is illustrated. Figures 8A to 8IA flowchart of a method that is an example of a computer system for reducing the visual prominence of virtual content relative to a three-dimensional environment, according to some embodiments. Figures 7A to 7E The user interface in Figures 8A to 8I illustrates the Figures 9A to 9E process in Figures 10A to 10F An example of a computer system for displaying an indication of a physical object in a three-dimensional environment, according to some embodiments. Figures 9A to 9E The user interface in Figures 10A to 10F illustrates the
[0026] The processes described below enhance the operability of the device and make the user-device interface more efficient through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow for the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage, which in turn reduces the heat emitted by the device, which is particularly important for wearable devices, where it can become uncomfortable for the user to wear the device if it generates too much heat within the operating parameters of the device components.
[0027] In addition, in a method where one or more of the steps described herein depend on one or more conditions being satisfied, it should be understood that the method can be repeated in multiple iterations such that, during the repetition process, all of the conditions that determine the steps in the method are satisfied in different iterations of the method. For example, if a method requires performing a first step (if a condition is satisfied) and a second step (if the condition is not satisfied), a person of ordinary skill in the art will know to repeat the stated steps until both the condition being satisfied and the condition not being satisfied (in no particular order) occur. Thus, a method described as having one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of corresponding one or more conditions and is thus able to determine whether the possible scenarios have been satisfied without explicitly repeating the steps of the method until all of the conditions that determine the steps in the method are satisfied. A person of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.
[0028] In some embodiments, as Figure 1 shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., household appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).
[0029] When describing XR experiences, various terms are used to distinctively refer to several related but different environments that a user can sense and / or with which the user can interact (e.g., interact using inputs detected by the computer system 101 that generates the XR experience, where these inputs cause the computer system generating the XR experience to generate audio, visual, and / or tactile feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:
[0030] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.
[0031] Extended reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic system. In XR, a subset of a person's physical movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in the physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in the XR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. In another example, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.
[0032] Examples of XR include virtual reality and mixed reality.
[0033] Virtual Reality: A virtual reality (VR) environment is an environment that is designed to be a simulated environment based entirely on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in the VR environment by way of a simulation of the person's presence within the computer-generated environment and / or by way of a simulation of a subset of the person's physical movements within the computer-generated environment.
[0034] Mixed Reality: Compared to a VR environment that is designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment is an environment that is designed to include, in addition to computer-generated sensory inputs (e.g., virtual objects), sensory inputs or representations thereof from the physical environment. On the virtual continuum, a mixed reality environment is any condition between a fully physical environment as one end and a virtual reality environment as the other end, but excluding these two ends. In some MR environments, the computer-generated sensory inputs can respond to changes in the sensory inputs from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or their representations). For example, the system can cause movements such that a virtual tree appears stationary relative to the physical ground.
[0035] Examples of mixed reality include augmented reality and augmented virtuality.
[0036] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that the person, using the system, perceives the virtual objects superimposed over the physical environment. Alternatively, the system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with the virtual objects and presents the combination on the opaque display. The person, using the system, indirectly views the physical environment through the images or video of the physical environment and perceives the virtual objects superimposed over the physical environment. As used herein, the video of the physical environment displayed on the opaque display is referred to as "passthrough video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that the person, using the system, perceives the virtual objects superimposed over the physical environment. An augmented reality environment is also a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, in providing passthrough video, the system can transform one or more sensor images to impose an alternative perspective (e.g., viewpoint) that is different from the perspective captured by the imaging sensor. As another example, a representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not a true version of the originally captured image. As yet another example, a representation of the physical environment can be transformed by graphically eliminating portions thereof or blurring portions thereof.
[0037] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but a person's face is a realistic reproduction from an image of a physical person. As another example, a virtual object can adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can adopt a shadow that conforms to the positioning of the sun in the physical environment.
[0038] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generation components. In some embodiments, the region defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the region defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and the viewport boundary typically move as the one or more display generation components move (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet device or a smart phone as the user's hand moves). The user's viewpoint determines the content visible in the viewport, the viewpoint typically specifying a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For a head-mounted device, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment and an immersive experience while the user is using the head-mounted device. For a handheld or stationary device, the viewpoint moves as the handheld or stationary device moves and / or as the user's positioning relative to the handheld or stationary device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For a device that includes a display generation component with virtual passthrough, the portions of the physical environment visible (e.g., displayed and / or projected) via the one or more display generation components are based on the field of view of one or more cameras in communication with the display generation component, the one or more cameras typically moving as the display generation component moves (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet device or a smart phone as the user's hand moves), because the user's viewpoint moves as the field of view of the one or more cameras moves (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual object are updated based on the movement of the user's viewpoint)).For a display generation component with optical see-through, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more partially or fully transparent portions of the display generation component) are based on the user's field of view through the partially or fully transparent portion of the display generation component (e.g., for a head-mounted device, it moves as the user's head moves, or for a handheld device such as a tablet device or a smart phone, it moves as the user's hand moves), because the user's viewpoint moves as the user's field of view through the partially or fully transparent portion of the display generation component moves (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0039] Viewpoint-locked virtual object: When a computer system displays a virtual object at the same position and / or orientation in the user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In an embodiment where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user looks straight ahead, the user's viewpoint is at least a part of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In an embodiment where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of the viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an embodiment where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object".
[0040] Environment-Locked Visual Objects: When a computer system displays a virtual object at a location and / or orientation within the user's field of view, the virtual object is environment-locked (alternatively, "world-locked"), where the location and / or orientation is based on a location and / or object within a three-dimensional environment (e.g., a physical or virtual environment), such that the virtual object is selected and / or anchored with reference to the location and / or object. As the user's field of view moves, the location and / or object within the environment relative to the user's field of view changes, which causes the environment-locked virtual object to be displayed at a different location and / or orientation within the user's field of view. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's field of view. When the user's field of view is shifted to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center within the user's field of view (e.g., the location of the tree within the user's field of view is offset), the environment-locked virtual object locked to the tree is displayed to the left of center within the user's field of view. In other words, the location and / or orientation at which the environment-locked virtual object is displayed within the user's field of view depends on the location and / or orientation of the location and / or object within the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to a fixed location and / or object within a physical environment) in order to determine the orientation at which the environment-locked virtual object is displayed within the user's field of view. The environment-locked virtual object may be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object), or may be locked to a movable portion of the environment (e.g., a vehicle, animal, person, or even a representation of a portion of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's field of view) such that the virtual object moves as the field of view or that portion of the environment moves to maintain a fixed relationship between the virtual object and that portion of the environment.
[0041] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits lazy follow behavior, which reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting lazy follow behavior, when a movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) that the virtual object is following is detected, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., the portion of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits lazy follow behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold amount of movement, such as moving 0 degrees to 5 degrees or moving 0 cm to 50 cm). For example, when the reference point (e.g., the portion of the environment to which the virtual object is locked or the viewpoint) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed orientation relative to a different viewpoint or portion of the environment than the reference point to which the virtual object is locked), and when the reference point (e.g., the portion of the environment to which the virtual object is locked or the viewpoint) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed orientation relative to a different viewpoint or portion of the environment than the reference point to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases above a threshold (e.g., the "lazy follow" threshold) because the virtual object is moved by the computer system to maintain a fixed or substantially fixed orientation relative to the reference point. In some embodiments, the virtual object maintaining a substantially fixed orientation relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the orientation of the reference point).
[0042] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, the head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). The head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system can have a transparent or translucent display instead of an opaque display. The transparent or translucent display can have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system can also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Below with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), input device 125, output device 155, one or more sensors in sensor 190, and / or one or more peripherals in peripheral 195, or shares the same physical housing or support structure with one or more of the foregoing devices.
[0043] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least the visual component of the XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with respect to Figure 3 In some embodiments, the functionality of controller 110 is provided by and / or in combination with display generation component 120.
[0044] According to some embodiments, display generation component 120 provides an XR experience to a user when the user is virtually and / or physically present within scene 105.
[0045] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet device) configured to present XR content, and the user holds the device having a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with XR content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).
[0046] Although relevant features of the operating environment 100 are shown in Figure 1 for the sake of brevity and to not obscure more relevant aspects of the example embodiments disclosed herein, various other features are not illustrated.
[0047] Figure 2FIG. 0 is a block diagram of an example of a controller 110 in accordance with some embodiments. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that, for the sake of brevity and in order not to obscure more relevant aspects of the embodiments disclosed herein, various other features are not illustrated. To that end, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0048] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.
[0049] The memory 220 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, the memory 220 includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. The memory 220 includes non-transitory computer-readable storage medium. In some embodiments, the memory 220 or the non-transitory computer-readable storage medium of the memory 220 stores the following programs, modules, and data structures or subsets thereof, including an optional operating system 230 and an XR experience module 240.
[0050] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., single XR experiences of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0051] In some embodiments, the data acquisition unit 241 is configured to obtain data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1 and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.
[0052] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track the positioning / location of at least the display generation component 120 relative to Figure 1 the scene 105, and optionally the positioning / location relative to one or more of the tracking input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions and heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to Figure 1 the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 244 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the XR content displayed via the display generation component 120. The eye tracking unit 243 is described in more detail below with respect to Figure 5 .
[0053] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally by one or more of the output device 155 and / or the peripheral device 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0054] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0055] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.
[0056] In addition, Figure 2 Rather, it is more of a functional description of the various features that may be present in a particular implementation, as opposed to the structural schematic of the embodiments described herein. As will be appreciated by those of ordinary skill in the art, items shown separately may be combined, and some items may be separated. For example, Figure 2 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are allocated therein, will vary depending on the particular implementation, and in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for the particular implementation.
[0057] Figure 3FIG. is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.
[0058] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0059] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffractive, reflective, polarized, holographic, and other waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.
[0060] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0061] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures or subsets thereof, including optionally operating system 330 and XR rendering module 340.
[0062] Operating system 330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR rendering module 340 includes data acquisition unit 342, XR rendering unit 344, XR mapping generation unit 346, and data transmission unit 348.
[0063] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) at least from Figure 1 controller 110. To this end, in various embodiments, data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0064] In some embodiments, XR rendering unit 344 is configured to present XR content via one or more XR displays 312. To this end, in various embodiments, XR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0065] In some embodiments, XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. To this end, in various embodiments, XR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0066] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0067] Although the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data sending unit 348 are shown as residing on a single device (e.g., Figure 1 the display generation component 120 of []), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in separate computing devices.
[0068] In addition, Figure 3 is more of a functional description of the various features that may be present in a particular embodiment, as opposed to the schematic diagrams of the embodiments described herein. As will be recognized by those of ordinary skill in the art, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the functional modules shown separately in [] may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are allocated therein, will vary depending on the specific implementation, and in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the specific implementation.
[0069] Figure 4 is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1 ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the positioning / position of one or more parts of the user's hand, and / or one or more parts of the user's hand relative to Figure 1Movement of the scene 105 (e.g., relative to a portion of the physical environment around the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0070] In some embodiments, the hand tracking device 140 includes an image sensor 404 that captures three-dimensional scene information of at least the hand 406 of a human user (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.). The image sensor 404 captures hand images at a sufficient resolution such that the fingers and their corresponding positions can be distinguished. The image sensor 404 typically captures images of other parts of the user's body and may also or possibly capture images of all parts of the body, and may have zoom capabilities or a dedicated sensor with increased magnification to capture images of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 or a portion of its field of view is positioned relative to the user or the user's environment in a manner that defines an interaction space in which hand movements captured by the image sensor are treated as inputs to the controller 110.
[0071] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and in addition, possibly color image data) to the controller 110, which extracts high-level information from the map data. The high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application drives the display generation component 120 accordingly. For example, the user can interact with software running on the controller 110 by moving the hand 406 and changing the pose of his hand.
[0072] In some embodiments, image sensor 404 projects a speckle pattern onto a scene that includes hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of points in the scene at a particular distance from image sensor 404 relative to a pre-determined reference plane. In the present disclosure, it is assumed that image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods, such as stereoscopy or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.
[0073] In some embodiments, hand tracking device 140 captures and processes a time series of depth maps that include the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software may match these descriptors to patch descriptors stored in database 408 based on a previous learning process in order to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.
[0074] The software may also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation function described herein may alternate with the motion tracking function such that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find changes in the pose that occur on the remaining frames. Pose, motion, and gesture information is provided to an application running on controller 110 via the API described above. The program may, for example, move and modify an image presented on display generation component 120 in response to the pose and / or gesture information, or perform other functions.
[0075] In some embodiments, the gesture includes an air gesture. An air gesture is detected without the user touching an input element (or independent of an input element that is part of a device, such as computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on the detected movement of a part of the user's body (such as the head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including the movement of the user's body relative to an absolute reference (such as the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (such as the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the hand), and / or the absolute movement of a part of the user's body (such as a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0076] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (such as a virtual or mixed reality environment) performed by the movement of the user's fingers relative to other fingers or parts of the user's hand. In some embodiments, an air gesture is detected without the user touching an input element that is part of a device (or independent of an input element that is part of a device) and is based on the detected movement of a part of the user's body through the air (including the movement of the user's body relative to an absolute reference (such as the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (such as the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the hand), and / or the absolute movement of a part of the user's body (such as a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0077] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides information to a computer system about which user interface element is the target of a user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or touchpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in a particular implementation involving an air gesture, for example, the input gesture is detected in combination with (e.g., concurrently with) the movement of the user's finger and / or hand towards the user interface element while the attention (e.g., gaze) towards the user interface element is detected to perform a pinch and / or tap input, as will be described in more detail below.
[0078] In some embodiments, an input gesture that points to a user interface object is performed with direct or indirect reference to the user interface object. For example, the user input is performed directly on the user interface object by performing an input at a location corresponding to the positioning of the user's hand relative to the positioning of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, when the user's attention (e.g., gaze) to the user interface object is detected, the input gesture is performed indirectly on the user interface object based on the fact that the positioning of the user's hand while the user performs the input gesture is not at the location corresponding to the positioning of the user interface object in the three-dimensional environment. For example, for a direct input gesture, the user can direct the user's input to the user interface object by initiating a gesture at or near a location corresponding to the display positioning of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm from the outer edge of the option or within the central part of the option, or within a distance between 0 cm and 5 cm measured from the outer edge or the central part of the option). For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the display positioning of the user interface object).
[0079] In some embodiments, according to some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment. For example, the pinch inputs and tap inputs described below are performed as air gestures.
[0080] In some embodiments, the pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of a hand to contact each other, i.e., optionally followed by an immediate break in contact with each other (e.g., within 0 seconds to 1 second). A long pinch gesture as an air gesture includes the movement of two or more fingers of a hand contacting each other for at least a threshold amount of time (e.g., at least 1 second) before detecting a break in contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) detected consecutively and immediately with each other (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks the contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).
[0081] In some embodiments, a pinching and dragging gesture as an air gesture includes a pinching gesture (e.g., a pinching gesture or a long pinching gesture) performed in combination with (e.g., following) a dragging input that changes the positioning of the user's hand from a first positioning (e.g., the starting positioning of the drag) to a second positioning (e.g., the ending positioning of the drag). In some embodiments, the user maintains the pinching gesture while performing the dragging input and releases the pinching gesture (e.g., opens two or more of their fingers) to end the dragging gesture (e.g., at the second positioning). In some embodiments, the pinching input and the dragging input are performed by the same hand (e.g., the user pinches two or more fingers together and moves the same hand to a second positioning in the air using the dragging gesture). In some embodiments, the pinching input is performed by the user's first hand and the dragging input is performed by the user's second hand (e.g., while the user continues the pinching input with the user's first hand, the user's second hand moves in the air from a first positioning to a second positioning). In some embodiments, an input gesture as an air gesture includes an input performed using both of the user's hands (e.g., a pinching and / or tapping input). For example, the input gesture includes two (e.g., or more) pinching inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinching gesture (e.g., a pinching input, a long pinching input, or a pinching and dragging input) is performed using the user's first hand, and a second pinching input is performed using the other hand (e.g., the second of the user's two hands) in combination with performing the pinching input using the first hand. In some embodiments, the movement between the user's two hands (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands)
[0082] In some embodiments, a tapping input performed as an air gesture (e.g., pointing to a user interface element) includes the movement of the user's finger towards the user interface element, the movement of the user's hand towards the user interface element (optionally, the user's finger extends towards the user interface element), the downward movement of the user's finger (e.g., mimicking a mouse click movement or a tap on a touch screen), or other predefined movements of the user's hand. In some embodiments, a tapping input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tapping gesture movement, which is the movement of the finger or hand away from the user's viewpoint and / or towards the object that is the target of the tapping input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tapping gesture (e.g., the end of the movement away from the user's viewpoint and / or towards the object that is the target of the tapping input, the reversal of the movement direction of the finger or hand, and / or the reversal of the acceleration direction of the movement of the finger or hand).
[0083] In some embodiments, the user's attention is determined to be directed to a portion of a three-dimensional environment based on detection of a gaze directed to the portion of the three-dimensional environment (optionally, without requiring additional conditions). In some embodiments, the user's attention is determined to be directed to a portion of a three-dimensional environment based on detection of a gaze directed to the portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold of the portion of the three-dimensional environment, such that the device determines that the user's attention is directed to the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0084] In some embodiments, the detection of the readiness state configuration of the user or a part of the user is detected by the computer system. The detection of the readiness state configuration of the hand is used by the computer system as an indication that the user may be about to use one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand to interact with the computer system. For example, based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to prepare for a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm faces away from the user), based on whether the hand is in a predetermined orientation relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user above the user's waist and below the user's head or moving away from the user's body or legs) to determine the readiness state of the hand. In some embodiments, the readiness state is used to determine whether interactive elements of the user interface respond to attention (e.g., gaze) inputs.
[0085] In scenarios where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used in place of the positioning and / or movement of one or more hands in the corresponding air gesture. In scenarios where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures. User input can be detected using controls contained within the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or change in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, where user input using the controls contained within the hardware input device is used in place of hand and / or finger gestures such as an air tap or air pinch in the corresponding air gesture. For example, a selection input described as being performed using an air tap or air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag can alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input after the movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space). Similarly, two-handed input including movement of the hands relative to each other can be performed using one air gesture and a hardware input device in a hand not performing the air gesture, two hardware input devices held in different hands, or two air gestures performed using different hands and / or various combinations of inputs detected by the one or more hardware input devices described above.
[0086] In some embodiments, the software can be downloaded electronically, for example, over a network, to the controller 110, or can alternatively be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is likewise stored in a memory associated with the controller 110. Alternatively or additionally, some or all of the described functions of the computer can be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although in Figure 4The controller 110 is shown, but for example, some or all of the processing functions of the controller, as a unit separate from the image sensor 404, may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device (such as a game console or a media player). The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device to be controlled by the sensor output.
[0087] Figure 4 Also shown is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels having corresponding depth values. The pixels 412 corresponding to the hand 406 have been segmented from the background and the wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from the image sensor 404), where the gray shading becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment the components (i.e., a group of adjacent pixels) of the image having human hand characteristics. These characteristics may include, for example, the overall size, shape, and motion from frame to frame in a sequence of depth maps.
[0088] Figure 4 Also illustrated schematically is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406. In Figure 4 it, the hand skeleton 414 is superimposed on the hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, the center of the palm, the end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the positions and movements of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.
[0089] Figure 5 An example embodiment of an eye tracking device 130 ( Figure 1 ) is illustrated. In some embodiments, the eye tracking device 130 is composed of an eye tracking unit 243 ( Figure 2)Control is used to track the positioning and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.
[0090] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, a head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and virtual objects are displayed on the transparent or translucent display, through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or projected as a hologram such that an individual uses the system to observe the virtual objects superimposed over the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0091] As Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be pointed at the user's eyes to receive the IR or NIR light directly reflected from the eyes by the light source, or alternatively can be pointed at a "hot" mirror located between the user's eyes and the display panel, which reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 to 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked individually by corresponding eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by corresponding eye tracking cameras and illumination sources.
[0092] In some embodiments, a device-specific calibration process is used to calibrate the eye tracking device 130 to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, cameras, hot mirrors (if any), eye lenses, and display screens. The device-specific calibration process can be performed at the factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include an estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, interocular distance, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's fixation point relative to the display.
[0093] As Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 may be pointed at a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass through) (e.g., as shown in the top portion of Figure 5 ), or alternatively may be pointed at the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the bottom portion of Figure 5 ).
[0094] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for a left display panel and a right display panel) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash assist method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction in which the user is currently looking.
[0095] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined according to the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content in the view at least in part based on the user's current gaze direction. As another example, the controller may display specific virtual content in the view at least in part based on the user's current gaze direction. As another example use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment of the XR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 so that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus so that a nearby object that the user is looking at appears at the correct distance.
[0096] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lens 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each of the lenses, as Figure 5 shown. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0097] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0098] As Figure 5 illustrated, embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience to the user.
[0099] Figure 6 illustrates a flash-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., an eye tracking device 130 as Figure 1 and Figure 5 illustrated). The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses the previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.
[0100] As Figure 6 shown, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be input into the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.
[0101] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0102] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of flashes for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, then at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking state is set to "yes" (if it is not already "yes"), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.
[0103] Figure 6 It is intended to be used as an example of an eye tracking technique that can be used for a particular specific implementation. As will be recognized by those of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing an XR experience to a user, other eye tracking techniques that currently exist or are developed in the future can be used to replace the flash-assisted eye tracking technique described herein or used in combination with the flash-assisted eye tracking technique.
[0104] In some embodiments, a captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed over a representation of the real-world environment 602.
[0105] Accordingly, the description herein describes some implementations of a three-dimensional environment (e.g., an XR environment) that includes a representation of real-world objects and a representation of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and a display of a computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, where the three-dimensional environment is based on the physical environment captured by one or more sensors of the computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in the three-dimensional environment at corresponding locations that have corresponding positions in the real world such that the virtual objects appear as if they exist in the real world (e.g., the physical environment). For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some implementations, the corresponding locations in the three-dimensional environment have corresponding positions in the physical environment. Thus, when the computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., a location at or near the user's hand or a location at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a location in the three-dimensional environment corresponding to the location where the virtual object would be displayed in the physical environment if the virtual object were a real object at that specific location).
[0106] In some implementations, real-world objects that exist in the physical environment and are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that only exist in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.
[0107] In a three-dimensional environment (e.g., a real environment, a virtual environment, or an environment that includes a mixture of real and virtual objects), an object is sometimes said to have depth or simulated depth, or an object is said to be visible, displayed, or placed at different depths. In this context, depth refers to a dimension other than height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to the position or viewpoint of a user, in which case the depth dimension varies based on the position of the user and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the position of the user relative to the surface of the environment (e.g., the surface of the floor or ground of the environment), an object that is further from the user along a line extending parallel to the surface is considered to have greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the position of the user and is parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user is positioned at the center of a cylinder that extends from the user's head towards the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., the direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), an object that is further from the user's viewpoint along a line extending parallel to the user's viewpoint is considered to have greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from a line extending from the user's viewpoint and parallel to the direction of the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere that extends outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which an application and / or system content is displayed), where the user interface container has a height and / or width, and depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments where depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the depth dimension for the container extends away from the user or the user's viewpoint), the height and / or width of the container is generally orthogonal or substantially orthogonal to a line that extends from the user's position (e.g., the user's viewpoint or the position of the user) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the positioning of the object along the depth dimension for the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions that extend in different directions and / or away from the user or the user's viewpoint from different starting points).In some embodiments, when depth is defined relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewing point changes (e.g., or when multiple different observers are viewing the same container in a three-dimensional environment, such as during a face-to-face collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content including the container). In some embodiments, for a curved container (e.g., including a container with a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, z-spacing (e.g., the spacing between two objects in the depth dimension), z-height (e.g., the distance of an object from another object in the depth dimension), z-positioning (e.g., the positioning of an object in the depth dimension), z-depth (e.g., the positioning of an object in the depth dimension), or an analog z-dimension (e.g., depth used as a dimension of an object, a dimension of the environment, a direction in space, and / or a direction in an analog space) are used to refer to the concept of depth as described above.
[0108] Similarly, the user optionally can interact with virtual objects in a three-dimensional environment using one or both hands as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of the computer system optionally capture one or more hands of the user and display a representation of the user's hand(s) in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, the user's hand(s) are visible via the display generation component, via the ability to see the physical environment through the user interface, due to the transparency / translucency of a portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / translucent surface or onto the user's eyes or into the user's field of view. Thus, in some embodiments, the user's hands are displayed at corresponding positions in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if the virtual objects were physical objects in a physical environment. In some embodiments, the computer system can update the display of the representation of the user's hand(s) in the three-dimensional environment in conjunction with the movement of the user's hand(s) in the physical environment.
[0109] In some of the embodiments described below, the computer system is optionally able to determine an “effective” distance between a physical object in the physical world and a virtual object in a three-dimensional environment, e.g., for determining whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, holding, etc. the virtual object or is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger of the hand pressing a virtual button, a hand of the user grasping a virtual vase, the hands of the user brought together and pinching / holding the user interface of an application, and two fingers performing any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a particular location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a particular corresponding location in the three-dimensional environment (e.g., the location at which the hand would be displayed in the three-dimensional environment if the hand were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing the location in the physical world (e.g., rather than comparing the location in the three-dimensional environment). For example, when determining the distance between one or more hands of the user and a virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., the location at which the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or to map the position of the virtual object to the physical environment.
[0110] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if the user's gaze is directed at a particular location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed at the virtual object. Similarly, the computer system is optionally able to determine the direction in the physical environment that the stylus is directed based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines the corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment where the stylus is directed, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.
[0111] Similarly, the embodiments described herein may refer to the location of the user (e.g., the user of the computer system) in the three-dimensional environment and / or the location of the computer system in the three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system serves as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to the corresponding location in the three-dimensional environment. For example, the location of the computer system will be the location in the physical environment (and its corresponding location in the three-dimensional environment) from which the user would see the objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other) as the objects that are displayed in the three-dimensional environment by the display generation component of the computer system or are visible in the three-dimensional environment via the display generation component, if the user were standing at that location and facing the corresponding portion of the physical environment visible via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed at the same location in the physical environment as the location of these virtual objects in the three-dimensional environment, and physical objects having the same size and orientation in the physical environment as when in the three-dimensional environment), the location of the computer system and / or the user is the location from which the user would see the virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation component of the computer system.
[0112] In this disclosure, various input methods are described in relation to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the input device or input method described in relation to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the output device or output method described in relation to the other example. Similarly, various methods are described in relation to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the methods described in relation to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples without having to list exhaustively all the features of the embodiments in the description of each example embodiment.
[0113] User interface and associated processes
[0114] Attention is now turned to embodiments of a user interface (“UI”) and associated processes that can be implemented on a computer system (such as a portable multifunctional device or a head-mounted device) having a display generation component, one or more input devices, and optionally one or more cameras.
[0115] Figures 7A to 7E An example of a computer system that reduces the visual prominence of virtual content with respect to a three-dimensional environment in accordance with some embodiments is illustrated.
[0116] Figure 7A An example of a three-dimensional environment 702 that is visible via a display generation component (e.g., Figure 1 display generation component 120) of computer system 101 is illustrated. The three-dimensional environment 702 is visible from the user's viewpoint 726, which is illustrated in a top view (e.g., facing the back wall of the physical environment in which computer system 101 is located). As described above with reference to Figures 1 to 6 computer system 101 optionally includes a display generation component (e.g., a touchscreen, an HMD or other wearable computing device, a display, a projector, etc.) and a plurality of image sensors (e.g., Figure 3The image sensor 314). The image sensor optionally includes one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a part of the user (e.g., one or more hands of the user) when the user interacts with the computer system 101. In some embodiments, the user interfaces illustrated and described below may also be implemented on a head-mounted display that includes a display generation component for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting the movement of the physical environment and / or the user's hand (e.g., external sensors facing away from the user) and / or sensors for detecting the user's gaze (e.g., internal sensors facing inward toward the user's face).
[0117] As Figure 7A shown, the computer system 101 captures one or more images of the physical environment around the computer system 101 (e.g., the operating environment 100), including one or more objects in the physical environment around the computer system 101. In some embodiments, the computer system 101 displays a representation of the physical environment in a three-dimensional environment 702, and / or the physical environment is visible in the three-dimensional environment 702 via the display generation component 120. For example, the three-dimensional environment 702 visible via the display generation component 120 includes a representation of the physical floor of the room in which the computer system 101 is located and the back and side walls. For example, and as Figure 7A shown, the physical environment around the computer system 101 includes a couch 724 (shown in a top view), which is visible from the Figure 7A user's viewpoint 726 in
[0118] In Figure 7A the three-dimensional environment 702 includes a user interface 704 and a user interface 706. In Figure 7A the user interfaces 704 and 706 are two-dimensional objects. It should be understood that the examples of the present disclosure optionally apply equally to three-dimensional objects. The user interfaces 704 and 706 are optionally one or more of the following: a user interface of an application (e.g., a messaging user interface and / or a content browsing user interface), a three-dimensional object (e.g., a virtual clock, a virtual ball, and / or a virtual car), or any other element displayed by the computer system 101 that is not included in the physical environment of the computer system 101. In Figures 7A to 7EIn [the context], the user interface 704 is a content playback user interface (e.g., a user interface through which a computer system 101 displays video and / or presents audio content), and the user interface 706 is a system user interface (e.g., a user interface including one or more controls that can be selected to control the computer system 101 and / or one or more functions of a system virtual environment displayed within the three-dimensional environment 702). The reference method 800 provides one or more details of the user interface 704 and the user interface 706.
[0119] In Figure 7A [the context], the computer system 101 is also presenting visual environmental effects to the three-dimensional environment 702, such as dimming, desaturating, or otherwise reducing the visual prominence of one or more portions of the physical environment visible via the display generation component 120. In some embodiments, the computer system 101 additionally or alternatively presents audio environmental effects that at least partially obscure one or more portions of the physical audio and / or virtual audio that would otherwise be audible from the physical environment.
[0120] In some embodiments, the computer system 101 presents the three-dimensional environment 702 in a manner that occludes or otherwise reduces the visibility of one or more portions of the physical environment surrounding the computer system 101. For example, Figure 7A the user interface 704 in [the context] occludes and / or obscures a portion of the back wall and the floor in the physical environment, and Figure 7A the user interface 706 in [the context] occludes and / or obscures a portion of the back wall and the floor in the physical environment. As another example, in Figure 7A [the context], the audio environmental effects and / or visual environmental effects presented by the computer system 101 are also occluding and / or obscuring a portion of the couch 724 and / or the audio components in the physical environment.
[0121] In some embodiments, and as discussed in more detail below, the computer system 101 reduces the visual prominence of one or more virtual portions of the three-dimensional environment 702 based on one or more characteristics of objects in the physical environment to allow those physical objects to become more visible within the three-dimensional environment 702. In some embodiments, the computer system 101 performs the above operations on non-system user interfaces (e.g., the user interface 704), but does not perform the above operations on system user interfaces (e.g., the user interface 706), as will be discussed in more detail below.
[0122] In Figure 7A [the context], a person 708 and a person 710 (two people other than the user of the computer system 101) are present within the physical environment surrounding the computer system 101. As Figure 7AAs shown, persons 708 and 710 are outside the field of view of the three-dimensional environment 702 visible via the display generation component 120 and are not visible within the three-dimensional environment 702.
[0123] From Figures 7A to 7B , persons 708 and 710 have moved within the physical environment around the computer system 101 and are now partially behind the user interfaces 704 and 706, respectively. As Figure 7B shown, within the three-dimensional environment 702, the user interface 704 at least partially obscures person 708, and the user interface 706 at least partially obscures person 710. This obscuring is optionally due to persons 708 and 710 not exhibiting one or more characteristics corresponding to the computer system 101 designating persons 708 and 710 as participating in a significant social interaction with the user of the computer system 101, as will be described in more detail later. Additionally, and because persons 708 and 710 are not designated as participating in a significant social interaction, and as Figure 7B shown, the audio and / or audio effects of persons 708 and 710 (e.g., the audio corresponding to the speech from person 708 or person 710) are blocked. Furthermore, it should be noted that because persons 708 and 710 have been determined not to exhibit characteristics of a significant social interaction, the audio environment effects and / or visual environment effects of the user interfaces 704 and 706 are Figure 7B the same in Figure 7A as in Figure 7B shown, the unobscured portions of person 708 by the user interface 704 (e.g., the feet of person 708) and the unobscured portions of person 710 by the user interface 706 (e.g., the feet of person 710) are visible within the three-dimensional environment 702 to which visual environment effects (e.g., dimming) are applied.
[0124] From Figures 7B to 7C , persons 708 and 710 exhibit one or more characteristics corresponding to the computer system 101 designating person 708 and / or person 710 as participating in a significant social interaction with the user of the computer system 101. The reference method 800 provides details regarding the various criteria used by the computer system 101 to designate such significant social interactions. For example, in Figure 7C shown, persons 708 and 710 have turned and / or are talking to the user of the computer system 101. Therefore, person 708 is eligible to penetrate the user interface 704, as Figure 7C shown. However, person 710 is not eligible to penetrate the user interface 706 because the user interface 706 is the system user interface. Therefore, and as Figure 7C shown and discussed in more detail below, person 708 penetrates the user interface 704 and person 710 does not penetrate the user interface 706.
[0125] For example, and asFigure 7C As shown, the action of person 708 turning and talking to the user of computer system 101 is recognized by computer system 101 as a significant social interaction, and computer system 101 takes actions to increase the visibility and / or audio effects of person 708. For example, and as Figure 7C shown, computer system 101 reduces (e.g., by 5%, 10%, 15%, or 25%) the visual prominence of portions of the three-dimensional environment 702 that include and / or surround person 708 to allow person 708 to be more visible (e.g., by 5%, 10%, 15%, or 25%) within the three-dimensional environment 702. The reduction in visual prominence optionally includes decreasing the magnitude or amount of the opacity of the environmental effects and / or user interface 704, reducing the brightness or saturation of the environmental effects and / or user interface 704, and / or reducing the number or density of particles in the particle effects. In some embodiments, the environmental effects are optionally reduced in the region and / or volume that includes and / or surrounds person 708. For example, in some embodiments, the environmental effects are reduced in the region and / or volume between person 708 and the viewing point 726 (e.g., the projection of person 708 towards the viewing point 726). In this way, the user's view of person 708 (e.g., the portion of the three-dimensional environment 702 between the viewing point 726 and person 708) is less occluded compared to before penetration.
[0126] In addition, and as Figure 7C shown, computer system 101 penetrates the user interface 704 such that person 708 is visible through the user interface 704. It should be noted that although the physical environment is optionally shown through the penetration of the user interface 704, person 708 occludes the physical environment directly behind person 708. In addition, and as Figure 7C shown, the non-penetrated portion of the user interface 704 maintains its previous visual prominence such that the user can still view portions of the user interface 704. In addition, and as Figure 7C shown, the portions of person 708 that are not obscured by the user interface 704 (e.g., the feet of person 708) are more visible in the three-dimensional environment because the visual prominence in the region around person 708 is reduced. In addition, and as Figure 7C shown, the portions of person 710 that are not obscured by the user interface 706 (e.g., the feet of person 710) are visible in the three-dimensional environment 702 to which visual environmental effects (e.g., dimming) are applied. Additionally, and as Figure 7C shown, the visual prominence of the region and / or volume outside the projection of person 708 (e.g., the three-dimensional environment 702 that includes the non-penetrated portion of the user interface 704) is not reduced. In addition, and as Figure 7CAs shown, the computer system 101 reduces the audio effects of the three-dimensional environment 702 to allow audio to be heard and / or generated from the person 708. For example, the audio effects from the person 708 optionally include the person 708 talking to the user of the computer system 101 and / or playing audio from a device (e.g., a mobile phone, a tablet device, and / or a stereo system).
[0127] As Figure 7C shown, the action of the person 710 turning and talking to the user of the computer system 101 is recognized by the computer system 101 as a significant social interaction. However, the computer system 101 does not take actions to increase the visibility of the person 710 and / or the audio effects from that person because the user interface 706 is the system user interface. Thus, the computer system 101 maintains the visual prominence of the portion of the three-dimensional environment 702 surrounding the person 710 and / or corresponding to the projection of the person 710 towards the viewing point 726 (e.g., including the user interface 706). For example, in some embodiments, the environmental effects are maintained in the region and / or volume between the person 710 and the viewing point 726. Additionally, and as Figure 7C shown, the environmental effects are maintained in the region and / or volume surrounding the person 710. Additionally, the audio effects and / or environmental effects for the user interface 706 are maintained, and the audio effects for the person 710 do not penetrate.
[0128] In some embodiments, the user of the computer system 101 can provide a variable input (e.g., in the form of a slider input) for controlling the time and / or degree at which penetration occurs. For example, the variable input optionally controls the amount of interaction of the person 708 and / or 710 that needs to occur to be classified as a significant social interaction. Additionally or alternatively, the variable input optionally controls the amount of audio penetration and / or visual penetration of the three-dimensional environment 702 that occurs once a significant social interaction is determined. The reference method 800 provides details regarding the various criteria used by the computer system 101 for specifying such significant social interactions. In Figure 7C this case, the user of the computer system 101 has optionally set the above slider to a specific amount of penetration and / or a specific penetration threshold. For example, from Figures 7C to 7D this, the user of the computer system 101 has optionally set the above slider to a higher amount of penetration and / or a lower penetration threshold. Thus, and as Figure 7D shown, since the action of the person 708 turning and talking to the user of the computer system 101 is recognized by the computer system 101 as a significant social interaction, the computer system 101 takes actions to increase the visibility of the person 708 and / or the audio effects from that person according to the amount of penetration set in the above slider. For example, and as Figure 7CAs shown, computer system 101 reduces (e.g., by 25%, 50%, 75%, or 90%) the visual prominence of portions of the three-dimensional environment 702 that include and / or surround person 708 to allow person 708 to be more visible (e.g., by 25%, 50%, 75%, or 90%) within the three-dimensional environment 702. The reduction in visual prominence optionally includes reducing the magnitude or amount of the opacity of the environmental effects and / or user interface 704, reducing the brightness or saturation of the environmental effects and / or user interface 704, and / or reducing the number or density of particles in the particle effects. In some embodiments, the environmental effects are optionally reduced (e.g., by 25%, 50%, 75%, or 90%) within a region and / or volume that includes and / or surrounds person 708. For example, in some embodiments, the environmental effects are reduced (e.g., by 25%, 50%, 75%, or 90%) within the region and / or volume between person 708 and the viewing point 726 (e.g., the projection of person 708 towards the viewing point 726). In this way, the user's view of person 708 (e.g., the portion of the three-dimensional environment 702 between the viewing point 726 and person 708) is less occluded than before penetration.
[0129] In addition, and as Figure 7D shown, computer system 101 penetrates user interface 704 such that person 708 is visible through user interface 704. It should be noted that although the physical environment is optionally shown through the penetration of user interface 704, person 708 occludes the physical environment directly behind person 708. In addition, and as Figure 7D shown, the non-penetrated portions of user interface 704 maintain their previous visual prominence such that the user can still view portions of user interface 704. In addition, and as Figure 7D shown, the portions of person 708 that are not obscured by user interface 704 (e.g., the feet of person 708) are more visible within the three-dimensional environment 702 because the visual prominence of user interface 704 and / or the environmental effects within the region surrounding person 708 are reduced (e.g., by 25%, 50%, 75%, or 90%). In addition, and as Figure 7D shown, the portions of person 710 that are not obscured by user interface 706 (e.g., the feet of person 710) are visible within the three-dimensional environment 702 to which visual environmental effects (e.g., dimming) are applied. Additionally, and as Figure 7D shown, the visual prominence of the region and / or volume outside the projection of person 708 (e.g., the three-dimensional environment 702 including the non-penetrated portions of user interface 704) is not reduced. In addition, and as Figure 7DAs shown, the computer system 101 further reduces the audio effects of the three-dimensional environment 702 to allow additional (e.g., 25%, 50%, 75%, or 90%) hearing and / or generation of audio from the person 708. For example, the audio effects from the person 708 optionally are the person 708 talking to the user of the computer system 101 and / or playing audio from a device (e.g., a mobile phone, a tablet device, and / or a stereo system).
[0130] As Figure 7D shown, the action of the person 710 turning and talking to the user of the computer system 101 is still recognized by the computer system 101 as a significant social interaction. However, the computer system 101 still does not take an action to increase the visibility of the person 710 and / or the audio effects from that person because the user interface 706 is the system user interface. Thus, the computer system 101 maintains the visual prominence of the portion of the three-dimensional environment 702 surrounding the person 710. For example, in some embodiments, the environmental effects are maintained in the area and / or volume between the person 710 and the viewing point 726. Additionally, and as Figure 7D shown, the environmental effects are maintained in the area and / or volume surrounding the person 710. Further, the audio effects and / or environmental effects for the user interface 706 are maintained, and the audio for the person 710 does not penetrate.
[0131] In some embodiments, the amount of penetration optionally depends on the position of the object relative to the content being penetrated (e.g., the user interface). For example, from Figures 7D to 7E , the person 708 has moved closer to the edge of the user interface 704. In some embodiments, the computer system 101 optionally fully penetrates an object whose action is recognized as a significant social interaction and is positioned towards the edge of the content (e.g., the user interface 704), such as within a threshold distance (e.g., 0.1 meter, 0.5 meter, 1 meter, 3 meters, 5 meters, 10 meters, or 20 meters) of the edge of the content. In some embodiments, and as Figure 7E shown, as the physical object gets closer to the edge of the user interface (e.g., the user interface 704), the computer system 101 increases the visibility of the physical object to be more visible in the three-dimensional environment 702. For example, and as Figure 7EAs shown, the computer system 101 reduces (e.g., reduces by 25%, 50%, 75%, or 100%) the visual prominence of portions of the three-dimensional environment 702 that include and / or surround the person 708 to allow the person 708 to be visible (e.g., 25%, 50%, 75%, or 100%) within the three-dimensional environment 702. The reduction of visual prominence optionally includes removing environmental effects and / or opacity of the user interface 704, removing brightness or saturation of the environmental effects and / or the user interface 704, and / or removing the number or density of particles in the particle effects. In some embodiments, the environmental effects are optionally removed in the region and / or volume that includes and / or surrounds the person 708. For example, in some embodiments, the environmental effects are removed in the region and / or volume (e.g., the projection of the person towards the viewing point 726) between the person 708 and the viewing point 726. In this way, the user's view of the person 708 (e.g., including the portion of the three-dimensional environment 702 between the viewing point 726 and the person 708) is less occluded than before the penetration.
[0132] In addition, and as Figure 7E shown, the computer system 101 penetrates the user interface 704 such that the person 708 is fully visible through the user interface 704. It should be noted that although the physical environment is optionally shown by the penetration of the user interface 704, the person 708 occludes the physical environment directly behind the person 708. In addition, and as Figure 7E shown, the non-penetrated portions of the user interface 704 maintain their previous visual prominence such that the user can still view portions of the user interface 704. In addition, and as Figure 7E shown, the computer system 101 reduces (e.g., reduces by 50%, 75%, 90%, or 100%) the audio effects of the three-dimensional environment 702 to allow audio from the person 708 to be heard.
[0133] Figures 8A to 8I is a flowchart of an exemplary method 800 for reducing the visual prominence of virtual content relative to a three-dimensional environment according to some embodiments. In some embodiments, the method 800 is executed at a computer system (e.g., Figure 1 the computer system 101 in, such as a tablet device, a smart phone, a wearable computer, or a head-mounted device), the computer system including a display generation component (e.g., Figure 1 、 Figure 3 and Figure 4a display generation component 120) (e.g., a head-up display, a monitor, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 800 is managed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., Figure 1 the controller 110) in ). Some operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0134] In some embodiments, method 800 is executed at a computer system (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., 314). For example, a mobile device (e.g., a tablet device, a smart phone, a media player, or a wearable device), a computer, or other electronic device. In some embodiments, the display generation component is a monitor integrated with the electronic device (optionally a touch screen monitor), an external monitor such as a monitor, a projector, a television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users, etc. In some embodiments, the one or more input devices include those capable of receiving user input (e.g., capturing user input, detecting user input, etc.) and sending information associated with the user input to the computer system. Examples of input devices include a touch screen, a mouse (e.g., external), a touchpad (optionally integrated or external), a touch panel (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the computer system), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye-tracking device, and / or a motion sensor (e.g., a hand-tracking device, a hand motion sensor), etc. In some embodiments, the computer system communicates with a hand-tracking device (e.g., one or more cameras, a depth sensor, a proximity sensor, a touch sensor (e.g., a touch screen, a touchpad)). In some embodiments, the hand-tracking device is a wearable device, such as a smart glove. In some embodiments, the hand-tracking device is a handheld input device, such as a remote control or a stylus.
[0135] In some embodiments, a three-dimensional environment is visible via the display generation component (e.g., generating, displaying a three-dimensional environment, or otherwise enabling it to be viewed by the computer system (e.g., an extended reality (XR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment, etc.)), such as Figure 7Avia the display generation component 120 in the three-dimensional environment 702. In some embodiments, the three-dimensional environment includes first virtual content (e.g., the first virtual content is a user interface of an application on a computer system (such as a content (e.g., movie, TV show, and / or music) playback application), and the user interface is displaying or otherwise presenting content; in some embodiments, the first virtual content is a two-dimensional or three-dimensional model of an object (such as a tent, building, or car) that is obscuring a first portion of the physical environment (such as Figure 7A the user interface 704 and / or user interface 706) in
[0136] In some embodiments, a first physical object (e.g., a person other than the user, an animal such as a pet, or any other physical object capable of moving) (802a) located at a first position in a first portion of the physical environment is detected via one or more input devices, e.g., Figure 7A the person 708 and / or person 710 in the physical environment in
[0137] In some embodiments, in response to detecting a first physical object in a first portion of the physical environment (e.g.,Figure 7C a person (708)(802b) within the physical environment, according to determining that the first physical object (e.g., 708) at the first position satisfies one or more first criteria (e.g., satisfaction of one or more first criteria corresponds to significant social interaction, and one or more criteria are discussed in more detail below such as with reference to steps 810 to 818) and the current value of the first user-defined setting is set to allow penetration of the physical object that satisfies one or more first criteria (e.g., the user can adjust at least one setting that controls whether the computer system adjusts the display of the first virtual content based on the presence of an object whose visibility is being obscured by the first virtual content and / or adjusts the degree of display of the first virtual content, as will be described in more detail below), reducing the visual prominence (802c) of the first virtual content relative to the three-dimensional environment (e.g., 702), e.g., the visual prominence of the user interface 704 relative to the three-dimensional environment 702 is reduced, as Figure 7C shown (e.g., including increasing the visibility of a first portion of the physical environment through the first virtual content). In some embodiments, the computer system completely stops the display of the first virtual content, such that the first physical object is completely visible through the first virtual content via the display generation component. In some embodiments, the computer system reduces or changes the opacity, brightness, color saturation, and / or other visual characteristics of the first virtual content to not completely reduce the visual prominence of the first virtual content relative to the three-dimensional environment, and increases the visibility of the first physical object through the first virtual content (e.g., increasing from 0% visible to 20%, 40%, 60%, or 80% visible, or increasing from 5% visible to 20%, 40%, 60%, or 80% visible). In some embodiments, the amount by which the computer system reduces the visual prominence of the first virtual content varies according to the current value of the first user-defined setting. In some embodiments, the visibility of the first physical object and / or the display of the first virtual content are controlled according to one or more steps in the method 1000 before the computer system reduces the visual prominence of the first virtual content relative to the three-dimensional environment. In some embodiments, the visual prominence of one or more portions of the first virtual content is reduced while maintaining the visual prominence of one or more other portions of the first virtual content, as will be described in more detail later. In some embodiments, the visual prominence of the entire first virtual content is reduced.
[0138] In some embodiments, based on determining that a first physical object (e.g., 708) at a first location satisfies one or more first criteria and the current value of a first user-defined setting is set to prevent penetration of physical objects that satisfy one or more first criteria, the visual prominence (802d) of a first virtual content is maintained relative to a three-dimensional environment (e.g., 702) (e.g., without increasing the visibility of a first portion of the physical environment through the first virtual content). For example, if a user of computer system 101 sets the current value of the first user-defined setting to prevent penetration of physical objects that satisfy one or more first criteria, and person 708 satisfies one or more first criteria, the visual prominence of the first virtual content is maintained relative to the three-dimensional environment 702 (e.g., as opposed to penetration as Figure 7C shown). In some embodiments, when the computer system maintains the visual prominence of the first virtual content relative to the three-dimensional environment, the visibility (or lack thereof) of the first physical object in the three-dimensional environment does not change. In some embodiments, the user is able to set controls or settings to prevent penetration of physical objects. Thus, even if the physical object satisfies one or more first criteria, the computer system maintains the visual prominence of the first virtual content relative to the three-dimensional environment. Selectively reducing the visual prominence of the virtual content allows the user to interact with the physical environment when needed while otherwise maintaining the display of the virtual content and reducing distraction to the user, which reduces input errors and makes interaction with the computer system more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0139] In some embodiments, in response to detecting a first physical object (e.g., 708) (804a) in a first portion of the physical environment, it is determined that the first physical object (e.g., 708) at the first location does not meet one or more first criteria (e.g., the characteristics of the first physical object or the actions of the first physical object do not amount to significant social interaction). Significant social interaction optionally refers to an interaction between the first physical object and a user of the computer system, where the user typically wants to hear / see such interaction. For example, if the first physical object is the user's spouse asking the user what they want for dinner, this is optionally an interaction that the user wants to see and / or hear, and such characteristics and / or actions of the first physical object optionally meet one or more first criteria. Alternatively, if the first physical object is the user's dog walking around the physical environment and playing with a toy, this is optionally an interaction that the user would not want to see and / or hear, and such characteristics and / or actions of the first physical object optionally do not meet one or more first criteria. For example, in some embodiments, one or more first criteria include criteria that are not met when the user's spouse is just walking around the physical environment. Additionally, in some embodiments, one or more first criteria include criteria that are not met when the user's spouse is talking to another person (other than the user of the computer system) or an animal in the physical environment. Conversely, in some embodiments, one or more first criteria include criteria that are met when the user's spouse turns and talks directly to the user), maintaining the visual prominence (804b) of the first virtual content (e.g., 704) relative to the three-dimensional environment (e.g., 702). For example, as Figure 7B shown, the person 708 does not meet one or more first criteria and the user interface 704 maintains its visual prominence relative to the three-dimensional environment 702. In some embodiments, and when the first physical object fails to meet one or more first criteria, when the computer system maintains the visual prominence of the first virtual content relative to the three-dimensional environment, the visibility (or lack thereof) of the first physical object in the three-dimensional environment does not change. Maintaining the visual prominence of the first virtual content when the first physical object does not meet one or more first criteria avoids interruption of the display of the first virtual content for physical objects that do not guarantee penetration, which reduces the need for additional input to return to the display of the virtual content, thereby reducing power usage and extending the battery life of battery-powered devices.
[0140] In some embodiments, when the current value of a first user-defined setting corresponds to the Do Not Disturb mode of a computer system, the current value of the first user-defined setting is set to prevent penetration of a physical object that meets one or more first criteria (e.g., 101). In some embodiments, a user of the computer system can set a control or setting to the Do Not Disturb mode to prevent penetration of physical objects. In the Do Not Disturb mode, even if a physical object meets one or more first criteria, the computer system maintains the visual prominence of the first virtual content relative to the three-dimensional environment. In some embodiments, during the Do Not Disturb mode, one or more notifications are suppressed (806) in response to detecting the generation of one or more notification events. In some embodiments, physical object penetration of virtual content is also suppressed during the Do Not Disturb mode. For example, if a user of computer system 101 sets the current value of the first user-defined setting to the Do Not Disturb mode and a person 710 meets one or more first criteria, the person 710 will not penetrate, such as in Figure 7C . In some embodiments, the Do Not Disturb mode also suppresses notifications of received messages, notifications, received calls, or other events that would cause the computer system to present a notification, whether audio, visual, and / or tactile. Controlling penetration in conjunction with the Do Not Disturb mode allows a user to choose to avoid interruptions when viewing the first virtual content and reduces the amount of input required to prevent penetration, which reduces power usage and extends the battery life of battery-powered devices.
[0141] In some embodiments, in response to detecting a first physical object (e.g., 708) in a first portion of the physical environment (808a), based on determining that the current value of a first user-defined setting corresponds to a dynamic setting for the first user-defined setting and meets one or more first criteria, and the satisfaction of the one or more first criteria is based on one or more characteristics of the first physical object (e.g., 708) (e.g., in the case where the first physical object is a person or an animal, based on whether and / or how the first physical object interacts with the user of the computer system, such as in a manner equivalent to significant social interaction, the one or more first criteria are optionally satisfied, such as more detailed discussion below with reference to steps 810 to 818), the visual prominence of the first virtual content (e.g., 704) is reduced (808b) relative to the three-dimensional environment (e.g., 702), such as Figure 7C shown (e.g., allowing the first physical object to penetrate the first virtual content). In some embodiments, the first user-defined setting is set to a dynamic setting. In such settings, determining whether penetration occurs for a physical object is optionally based on the determined attributes of the physical object, as will be described in more detail below.
[0142] In some embodiments, based on determining that the current value of a first user-defined setting corresponds to a dynamic setting for the first user-defined setting and does not meet one or more first criteria (e.g., because one or more of the characteristics of the first physical object as described below do not meet one or more first criteria), the visual prominence of a first virtual content (e.g., 704) is maintained with respect to a three-dimensional environment (e.g., 702) (e.g., the first physical object is not allowed to penetrate the first virtual content) (808c). For example, if a user of computer system 101 sets a first user-defined setting to a dynamic setting and it does not meet one or more criteria, the visual prominence of user interface 704 is maintained with respect to three-dimensional environment 702, such as Figure 7B shown. Using a dynamic setting to control penetration allows the computer system to appropriately determine whether an object penetrates while avoiding excessive or insufficient penetration, which reduces power usage and extends the battery life of battery-powered devices.
[0143] In some embodiments, the first physical object is a person other than the user of the computer system (e.g., 101) (e.g., the user's spouse, partner, child, friend, or just another person in the environment), and one or more first criteria include criteria that are met when the person's gaze is directed at the user of the computer system (e.g., 101) (810). For example, if person 708 gazes at the user of computer system 101, the criterion is met. In some embodiments, if a person gazes or directly looks at the user of the computer system, the action meets the criterion and corresponds to a significant social interaction. In some embodiments, the gaze must be detected within a pre-determined time period (e.g., 0.3 seconds, 0.5 seconds, 1 second, 3 seconds, 5 seconds, 10 seconds, 20 seconds, or 60 seconds) before the criterion is met. The time period is optionally set by the user or determined by the computer system. In some embodiments, if the person's gaze is not directed at the user of the computer system, the criterion is not met. Allowing penetration based on a person's gaze directed at the user of the computer system reduces the number of inputs required for the user to interact with the individual, which reduces input errors and makes the interaction with the computer system more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0144] In some embodiments, the first physical object is a person other than the user of the computer system (e.g., 101), and one or more of the first criteria include criteria that are satisfied when the person is oriented towards the user of the computer system (e.g., 101) (or when the person has changed the orientation of a part of their body such that they are oriented towards the user of the computer system). For example, if person 708 is oriented towards the user of computer system 101, the criteria are satisfied. In some embodiments, if a person turns towards the user of the computer system, the action satisfies the criteria and corresponds to a significant social interaction. In some embodiments, such turning towards the user is a partial or full turn towards the user of the computer system (e.g., turning towards the user such that the person's head, torso, shoulders, and / or other parts are within a threshold range of angles directly oriented towards the user (e.g., within 1 degree, 3 degrees, 5 degrees, 15 degrees, 30 degrees, 45 degrees, 60 degrees, 90 degrees, or 180 degrees)). In some embodiments, such turning of the user is required to continue for a set period of time (e.g., 0.3 seconds, 0.5 seconds, 1 second, 3 seconds, 5 seconds, 10 seconds, 20 seconds, or 60 seconds) before the criteria are satisfied. The period of time is optionally set by the user or determined by the computer system. In some embodiments, if a person is walking next to the user and / or not turning towards the user, the criteria are not satisfied. Allowing for penetration based on the orientation of the person relative to the user of the computer system reduces the number of inputs required for the user to interact with the individual, which reduces input errors and makes interaction with the computer system more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0145] In some embodiments, the first physical object is a person other than the user of the computer system (e.g., 101), and one or more first criteria include criteria (814) that are met when the person is talking to the user of the computer system (e.g., 101). For example, if person 708 is talking to the user of computer system 101, the criteria are met. In some embodiments, if the person is (directly) talking to the user, the action meets the criteria and corresponds to a significant social interaction. In some embodiments, if the person is talking to another person in the room but not (directly) to the user of the computer system, this does not meet the criteria and does not correspond to a significant social interaction. However, if the person first talks to another person in the room (which optionally does not meet the criteria), and subsequently starts (directly) talking to the user of the computer system, the criteria will optionally become met. In some embodiments, if the person is not talking to the user, the criteria are not met. In some embodiments, the computer system detects that the person is talking to the user of the computer system based on whether the person is oriented towards the user, whether the person's mouth is open and / or moving, and / or whether the computer system's microphone detects words spoken by the person. Allowing penetration based on the person talking to the user of the computer system reduces the number of inputs required for the user to interact with the person, which reduces input errors and makes interaction with the computer system more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0146] In some embodiments, the first physical object is a person other than the user of the computer system (e.g., 101), and one or more criteria include criteria (816) that are met when the person is a known person as specified by the user of the computer system (e.g., 101). For example, if person 708 is a known person specified by the user of computer system 101, the criteria are met. In some embodiments, if the person is identified as known to the user, the identification meets the criteria and corresponds to a significant social interaction. For example, the computer system optionally identifies the person as the user's partner or child. The identification is optionally performed by the computer system in various ways. For example, the identification is optionally performed by the computer system, etc., based on facial recognition. In some embodiments, the person known to the user is set by the user (e.g., the user registers the person in the computer system's facial recognition system) or determined by the system. In some embodiments, if the person is not known to the user and / or the computer system, the criteria are not met. Allowing penetration based on the computer system's identification of the person reduces the number of inputs required for the user to interact with the person, which reduces input errors and makes interaction with the computer system more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0147] In some embodiments, the first physical object is a person other than the user of the computer system (e.g., 101), and one or more first criteria include criteria (818) that are satisfied when the person is within a threshold distance of the user of the computer system (e.g., 101). For example, if person 708 is within the threshold distance of the user of computer system 101, the criteria are satisfied. In some embodiments, if the person is within the threshold distance of the user (e.g., 0.05 meters, 0.1 meters, 0.25 meters, 0.5 meters, 0.8 meters, 1 meter, 2 meters, 5 meters, 10 meters, or 20 meters), this satisfies the criteria and corresponds to significant social interaction. In some embodiments, the threshold is optionally achieved over time as the person walks around. For example, if a person enters a room and is not within the threshold distance of the user, the criteria will not be satisfied until the person walks closer to the user and falls within the threshold distance of the user. Additionally, if a person approaches the user and falls within the threshold distance (which will optionally satisfy the criteria), and then moves away from the user and falls outside the threshold distance, the criteria will optionally then become not satisfied. Allowing penetration based on the distance between the person and the user of the computer system reduces the number of inputs required for the user to interact with the person and also avoids collisions between the user and the person, which reduces input errors and makes interaction with the computer system more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0148] In some embodiments, the three-dimensional environment (e.g., 702) includes first virtual content (e.g., 704) and is displayed (820a) with environmental effects (such as Figure 7A and Figure 7E the environmental effects shown (e.g., a dimming or brightening effect of the three-dimensional environment, a virtual lighting effect on one or more portions of the three-dimensional environment, and / or a virtual particle effect in one or more portions of the three-dimensional environment). In some embodiments, in response to detecting a first physical object (e.g., 708) in a first portion of the physical environment (820b), the visual prominence of the environmental effects in the three-dimensional environment (e.g., 702) is reduced (820c) according to determining that the first physical object (e.g., 708) at the first position satisfies one or more first criteria and the current value of a first user-defined setting is set to allow penetration for physical objects that satisfy one or more first criteria. For example, if person 708 satisfies one or more first criteria and the current value of the first user-defined setting is set to allow penetration for physical objects that satisfy one or more first criteria, the visual prominence of the environmental effects in three-dimensional environment 702 is reduced, as Figure 7CAs shown. In some embodiments, when the first physical object meets one or more first criteria, as part of the physical object penetrating the first virtual content and / or the three-dimensional environment, the visibility of the environmental effects in the three-dimensional environment is reduced. In some embodiments, reducing the visual prominence of the environmental effects includes reducing the magnitude or amount of the environmental effects, reducing the brightness or saturation of the environmental effects, and / or reducing the number or density of the particles in the particle effects. Reducing the visual prominence of the environmental effects in the three-dimensional environment is beneficial to the visibility of the physical object during penetration, which reduces the need for additional input to display the visibility of the physical object, thereby reducing power consumption and extending the battery life of battery-powered devices.
[0149] In some embodiments, the environmental effects are displayed in a first portion of the three-dimensional environment (e.g., 702) that is within a threshold distance (e.g., 0.05 meters, 0.1 meters, 1 meter, 3 meters, 5 meters, or 10 meters) of the projection of the first physical object (e.g., 708) relative to the view point of a user (e.g., 726) of the computer system (e.g., 101), and the environmental effects are displayed in a second portion of the three-dimensional environment (e.g., 702) that is outside the threshold distance (822a) of the projection of the first physical object (e.g., 708) relative to the view point of the user (e.g., 726), such as Figure 7BThe environmental effects shown. In some embodiments, the environmental effects are displayed in multiple (or all) parts of the three-dimensional environment before the first physical object is detected in the first part of the physical environment. For example, when a user of a computer system is viewing a three-dimensional environment from a viewpoint, the environmental effects are applied to both the area around the first physical object and the area outside the first physical object. In some embodiments, the projection of the first physical object relative to the user's viewpoint corresponds to one or more parts of the three-dimensional environment that exist between the first physical object and the user's viewpoint (e.g., the logical extrusion of the first physical object along the user's line of sight from the viewpoint to the first physical object), and a threshold distance is measured perpendicular to such line of sight / viewpoint and extending from the outer boundary of such projection. In some embodiments, in response to detecting the first physical object (e.g., 708) (822b) in the first part of the physical environment, based on determining that the first physical object (e.g., 708) at the first location meets one or more first criteria and the current value of the first user-defined setting is set to allow penetration for physical objects that meet one or more first criteria, the visual prominence of the first virtual content is reduced relative to the three-dimensional environment (e.g., 702), and the visual prominence of the environmental effects in the first part of the three-dimensional environment is reduced, while maintaining the visual prominence of the environmental effects in the second part of the three-dimensional environment (822c). For example, if the person 708 meets one or more first criteria, the visual prominence of the environmental effects within the threshold distance of the projection of the person 708 is reduced, while maintaining the visual prominence of the environmental effects outside the threshold distance of the projection of the person 708, such as Figure 7C shown. In some embodiments, and in response to the first physical object meeting one or more first criteria, the visual prominence of the environmental effects within the threshold distance of the projection of the first physical object is reduced, while maintaining the visual prominence of the environmental effects outside the threshold distance of the projection of the first physical object. In this way, the area around the projection of the first physical object is penetrated, while not interrupting the display of other areas of the three-dimensional environment. Only penetrating the area around the projection of the first physical object allows the user to interact with the physical object, while still maintaining the display of the environmental effects outside the projection of the first physical object, which reduces the need for additional input to display virtual content, thereby reducing power consumption and extending the battery life of battery-powered devices.
[0150] In some embodiments, a first portion of first virtual content is displayed in a first portion of a three-dimensional environment that is within a threshold distance of a projection of a first physical object (e.g., 708) relative to a viewpoint of a user (e.g., 726) of a computer system, and a second portion of the first virtual content is displayed in a second portion of the three-dimensional environment that is outside the threshold distance of the projection of the first physical object (e.g., 708) relative to the user's (e.g., 726) viewpoint (824a) (e.g., as described above). In some embodiments, in response to detecting a first physical object (e.g., 708) in a first portion of the physical environment (824b), and based on determining that the first physical object (e.g., 708) at a first location satisfies one or more first criteria and a current value of a first user-defined setting is set to permit penetration of physical objects that satisfy the one or more first criteria, the visual prominence of the first portion of the first virtual content in the first portion of the three-dimensional environment (e.g., 702) is reduced while maintaining the visual prominence of the second portion of the first virtual content in the second portion of the three-dimensional environment (824c). For example, if person 708 satisfies one or more first criteria, the visual prominence of user interface 704 within the threshold distance of the projection of person 708 is reduced while maintaining the visual prominence of user interface 704 outside the threshold distance of the projection of person 708, such as Figure 7C as shown. In some embodiments, and in response to the first physical object satisfying one or more first criteria, the visual prominence of the first virtual content within the threshold distance of the projection of the first physical object is reduced while maintaining the visual prominence of the first virtual content outside the threshold distance of the projection of the first physical object. In this way, the region of the first virtual content surrounding the projection of the first physical object is penetrated while not interrupting the display of other regions of the first virtual content. Only penetrating the region of the first virtual content surrounding the projection of the first physical object allows the user to interact with the physical object while still maintaining the display of the first virtual content outside the projection of the first physical object, which reduces the need for additional input to display the first virtual content, thereby reducing power usage and extending the battery life of battery-powered devices.
[0151] In some embodiments, the three-dimensional environment (e.g., 702) further includes second virtual content, where a first physical object (e.g., 708) is located between the first virtual content (e.g., 704) and the second virtual content relative to the viewpoint of a user (e.g., 726) of the computer system (e.g., 101), and the second virtual content is farther from the viewpoint of the user (e.g., 726) than the first virtual content. (826a) In some embodiments, from the viewpoint of the user, the second virtual content is displayed behind a person who is displayed behind the first virtual content. In some embodiments, in response to detecting a first physical object (e.g., 708) in a first portion of the physical environment (826b), based on determining that the first physical object (e.g., 708) at a first location meets one or more first criteria and the current value of a first user-defined setting is set to allow penetration for physical objects that meet one or more first criteria, the visual prominence of the first virtual content is reduced relative to the three-dimensional environment (e.g., 702), while the visual prominence of the second virtual content is maintained relative to the three-dimensional environment (e.g., 702) (826c). For example, if a person 708 is standing between the second virtual content and the user interface 704, and the person 708 meets one or more first criteria, the person 708 penetrates the user interface 704, such as Figure 7C as shown, while maintaining the second virtual content. In some embodiments, the first physical object is optionally a person standing between two portions of the virtual content. In such a case, and in response to determining that the person meets one or more first criteria, the person penetrates the virtual content closest to the user of the computer system, while maintaining the farther virtual content. In some embodiments, the second virtual content is system content, such as a virtual environment as described below. Maintaining the visual prominence of the second virtual content allows the user to interact with the first physical object without interrupting the interaction with the second virtual content, which reduces the need for additional input to display the second virtual content, thereby reducing power usage and extending the battery life of battery-powered devices.
[0152] In some embodiments, the three-dimensional environment (e.g., 702) further includes second virtual content, where a first physical object (e.g., 708) is located between a first virtual content (e.g., 704) and the second virtual content relative to the viewpoint of a user (e.g., 726) of the computer system, and the second virtual content is farther from the user's (e.g., 726) viewpoint than the first virtual content (e.g., 704) (828a). In some embodiments, from the user's viewpoint, the second virtual content is displayed behind a person who is displayed behind the first virtual content. In some embodiments, in response to detecting a first physical object (e.g., 708) in a first portion of the physical environment (828b), based on determining that the first physical object (e.g., 708) at a first location meets one or more first criteria and the current value of a first user-defined setting is set to allow penetration for physical objects that meet one or more first criteria, the visual prominence of the first virtual content is reduced relative to the three-dimensional environment (e.g., 702), and the visual prominence of the second virtual content is reduced relative to the three-dimensional environment (e.g., 702) (828c). For example, if a person 708 is standing in the line of sight of a user of the computer system 101 and between the user interface 704 and the second virtual content, and the person 708 meets one or more criteria, the person 708 penetrates the user interface 704, such as Figure 7C shown, and the visual prominence of the second virtual content is also reduced. In some embodiments, the first physical object is optionally a person standing within the line of sight of a user of the computer system and between two portions of the virtual content. If the person meets the penetration criteria, the visual prominence of the two portions of the virtual content is reduced. Reducing the visual prominence of both the first virtual content and the second virtual content reduces distraction when interacting with the first physical object, which reduces the need for additional input for interacting with the first physical object, thereby reducing power usage and extending the battery life of battery-powered devices.
[0153] In some embodiments, in response to detecting a first physical object (e.g., 708) (830a) in a first portion of the physical environment, it is determined that the current value of a first user-defined setting is set to a first value. (In some embodiments, the first user-defined setting can be adjusted via a slider optionally displayed in the control user interface of the computer system. For example, the slider can optionally be adjusted up or down by the user to adjust one or more characteristics of the penetration, as will be described in more detail below.) And it is determined that one or more second criteria are met (e.g., the one or more second criteria are optionally criteria for controlling whether the first physical object penetrates the first virtual content, as described herein. In some embodiments, the one or more second criteria are the same as or different from the one or more first criteria), and the visual prominence of the first virtual content (e.g., 704) is reduced by a first amount (830b) relative to the three-dimensional environment (e.g., 702), e.g., the visual prominence of the user interface 704 is reduced by the first amount relative to Figure 7C the three-dimensional environment 702 shown. In some embodiments, the first amount is based on the first value set by the user. In some embodiments, if the one or more second criteria are not met, the visual prominence of the first virtual content is maintained. These and other details will be discussed in more detail below, such as with reference to steps 832 to 838.
[0154] In some embodiments, it is determined that the current value of the first user-defined setting is set to a second value different from the first value. (For example, the user has adjusted the first user-defined setting up or down to the second value via the slider. In some embodiments, a second value higher than the first value corresponds to a penetration that is easier to trigger and / or more penetrations.) And it is determined that one or more third criteria are met (in some embodiments, the one or more third criteria are optionally criteria for controlling whether the first physical object penetrates the first virtual content, as described herein, and are optionally the same as or different from the one or more second criteria, as will be described in more detail below), and the visual prominence of the first virtual content is reduced by a second amount (e.g., optionally the same as or different from the first amount) relative to the three-dimensional environment (e.g., 702), e.g., the visual prominence of the user interface 704 is reduced by the second amount relative to Figure 7D the three-dimensional environment 702 shown (830c). In some embodiments, the second amount is based on the second value set by the user. In some embodiments, if the one or more third criteria are not met, the visual prominence of the first virtual content is maintained. Allowing the user to adjust the penetration threshold provides the user with control over the penetration property, which reduces the need for additional input to achieve the desired penetration property, thereby reducing power usage and extending the battery life of battery-powered devices.
[0155] In some embodiments, one or more second criteria are different from one or more third criteria (832). In some embodiments, the one or more second criteria and the one or more third criteria respectively correspond to settings for determining whether penetration occurs based on a first value and a second value. For example, in some embodiments, if the user adjusts the slider to a low value (e.g., the first value is a low amount less than 50% of the value range of the slider), then the one or more second criteria will correspond to a relatively high (or relatively difficult to meet) setting for penetration. In one example, the duration required to meet the penetration criteria for a person's gaze will optionally be relatively high (e.g., the gaze length required by the second criteria is a high setting that exceeds 50% of the value range of the gaze length setting). In some embodiments, if the user adjusts the slider to a high value (e.g., the second value is a high amount greater than 50% of the value range of the slider), then the one or more third criteria will correspond to a relatively low (or relatively easy to meet) setting for penetration. For example, the duration required to meet the penetration criteria for a person's gaze will optionally be relatively low (e.g., the gaze length required by the second criteria is a low setting that is less than 50% of the value range of the gaze length setting). Thus, in some embodiments, the slider control adjusts the ease with which an object penetrates virtual content, where a lower slider value makes penetration more difficult and a higher slider value makes penetration less difficult. Allowing the user to adjust the settings for whether penetration occurs provides the user with control over the penetration conditions and better aligns the user's expectations of penetration with the operation of the computer system, which reduces the need for additional input to achieve the desired penetration conditions, thereby reducing power usage and extending the battery life of battery-powered devices.
[0156] In some embodiments, a second quantity is different from a first quantity (834), such as Figure 7C and Figure 7DAs shown. In some embodiments, the first amount and the second amount correspond to the amount of penetration (if any) that occurs for a given set of penetration trigger conditions (e.g., one or more second criteria are met, or one or more third criteria are met). For example, if the first amount is greater than the second amount, a greater area of the first virtual content is optionally penetrated compared to the second amount. In some embodiments, if the first amount is greater than the second amount, the penetrated area of the first virtual content has a higher transparency compared to the second amount. Additionally or alternatively, if the first amount is greater than the second amount, the opacity, brightness, or color saturation is optionally reduced (or more generally appropriately changed) by a greater amount compared to the second amount. One or more additional or alternative visual characteristics of the first virtual content are optionally changed (e.g., appropriately increased or decreased) by a greater amount such that the first virtual content obscures the physical environment behind it less for the first penetration amount compared to the second penetration amount. Allowing the user to adjust the amount of penetration that occurs enables the user to control how much the visual prominence of the virtual content is reduced when penetration occurs and better aligns the user's expectations of penetration with the operation of the computer system, which reduces the need for additional input to achieve the desired penetration, thereby reducing power usage and extending the battery life of battery-powered devices.
[0157] In some embodiments, user input corresponding to a request to set the current value of a first user-defined setting to a first value is detected via one or more input devices, such as an input that sets the user-defined setting to the first value, causing Figure 7C a response from the computer system 101 in (e.g., the user input includes adjusting a penetration slider to a position corresponding to the first value or the second value as described herein). In some embodiments, the user input is detected when the current value of the first user-defined setting corresponds to a dynamic setting for the first user-defined setting (836a). As described above, in some embodiments, the first user-defined setting is set to a dynamic setting such that determining whether penetration occurs for a physical object is optionally based on the determined attributes of the physical object. In some embodiments, in response to detecting the user input, the current value of the first user-defined setting is set to the first value (836b), and after setting the current value of the first user-defined setting to the first value (or the second value), the current value of the first user-defined setting is automatically adjusted to correspond to the dynamic setting (836c) based on determining that a period of time (e.g., 0.1 minute, 0.3 minute, 1 minute, 3 minute, 5 minute, 20 minute, 45 minute, 60 minute, 120 minute, or 240 minutes) has elapsed since receiving the user input, such as reverting to cause Figure 7DUser-defined settings for the response of a computer system. In some embodiments, and when the user overrides the dynamic mode for a first user-defined setting, the setting set by the user is only applied for a period of time. The period of time is optionally set by the user or determined by the computer system. In some embodiments, when the setting set by the user returns to the dynamic mode, the computer system will adjust the penetration accordingly. For example, if the dynamic setting requires a smaller amount of penetration, the penetration will be reduced compared to the user-defined value for the first user-defined setting. Alternatively, if the dynamic setting requires a larger amount of penetration, the penetration will be increased compared to the user-defined value for the first user-defined setting. Additionally, the dynamic setting can require the same amount of penetration as the user-defined value. In this case, the penetration will optionally remain the same. Automatically restoring the penetration setting to the dynamic mode after the period of time ensures that potentially unreasonable or unsafe penetration settings do not remain valid at the computer system, which reduces the need for additional input to restore the penetration setting, thereby reducing power usage and extending the battery life of battery-powered devices.
[0158] In some embodiments, the user input corresponds to a request to switch between a first mode and a second mode. In the first mode, the current value of the first user-defined setting is user-defined, and in the second mode, the current value of the first user-defined setting is dynamic (838). In some embodiments, a user of the computer system can switch the penetration setting between the dynamic mode and the slide zoom mode (e.g., user-defined mode). For example, a user of the computer system can optionally set the user-defined setting to a set amount at a first time and then decide that they want the computer system to operate in the dynamic mode. In such a scenario, the user can optionally switch the user-defined setting back to the dynamic mode. Subsequent user inputs can optionally continue to switch between the dynamic mode and the user-defined mode. This switching is optionally achieved through a slider, a button, etc. Enabling the user to switch between the penetration slider mode and the dynamic mode allows the user to easily adjust to new situations in the user's physical environment, which reduces the need for additional input to adjust to new situations, thereby reducing power usage and extending the battery life of battery-powered devices.
[0159] In some embodiments, one or more first criteria include criteria (840) that are satisfied when a first physical object (e.g., 708) is designated by the user as being allowed to penetrate, such as if Figure 7C the person 708 in Figure 7CResponse of computer system 101 shown to person 708. In some embodiments, if the first physical object is a specific object, the object meets the criteria. For example, the user of the computer system can optionally select or specify a particular physical object (e.g., a person or an animal / pet) that is allowed to penetrate. Determining the particular person or animal is optionally based on object and / or facial recognition (e.g., using manually specified rules and / or a machine learning model for identifying the object and / or face). In some embodiments, if the person entering the room is not the specified person or animal, the penetration criteria will not be met, even if the person (or physical object) would otherwise meet the conditions required for penetration. Allowing penetration based on the specified or selected object reduces the incidence of unwanted penetration, which reduces the need for additional input to enable penetration for the specified or selected object, thereby reducing power usage and extending the battery life of battery-powered devices.
[0160] In some embodiments, one or more first criteria include criteria (842) that are not met when the first physical object (e.g., 708) is specified by the user as not allowed to penetrate, such as if Figure 7B the person 708 in is specified as not allowed to penetrate, resulting in Figure 7B Response of computer system 101 shown to person 708. In some embodiments, if the first physical object is a specific object, the object does not meet the criteria. For example, the user of the computer system can optionally select or specify a particular physical object (e.g., a person or an animal / pet) that is not allowed to penetrate. Determining the particular person or animal is optionally based on object and / or facial recognition. In some embodiments, if the person entering the room is not the specified person or animal, the penetration criteria can be met. Preventing penetration based on the specified or selected object allows the user to maintain the display of virtual content and reduces distraction to the user, which reduces the need for additional input to return to the display of virtual content, thereby reducing power usage and extending the battery life of battery-powered devices.
[0161] In some embodiments, the first virtual content includes a boundary (e.g., the first virtual content is displayed in a rectangular region or other region of a three-dimensional environment having one or more boundaries such as a top boundary, a bottom boundary, a left boundary, and / or a right boundary), a first portion of the first virtual content, and a second portion of the first virtual content that is closer to the boundary than the first portion of the first virtual content (such as Figures 7A to 7Ethe outer boundary of the user interface 704 and portions of the user interface 704 that are closer to and farther from such boundaries). In some embodiments, the visual prominence of the first virtual content is reduced (844a) relative to the three-dimensional environment (e.g., 702). In some embodiments, the first virtual content has different portions, and in the case of a penetration, the visual prominence of these portions is reduced at different levels, as will be described in more detail below. The second portion of the first virtual content is optionally closer to the boundary than the first portion of the virtual content. For example, if the boundary is the right boundary of the first virtual content, the second portion of the first virtual content is the portion of the virtual content adjacent to the right boundary, and the first portion of the first virtual content is to the left of the second portion, such as the central portion of the first virtual content. In some embodiments, reducing includes reducing the visual prominence of the first portion of the first virtual content by a first amount (844b) relative to the three-dimensional environment (e.g., 702), such as Figure 7D the reduction in the visual prominence of the portion of the user interface 704 shown. In some embodiments, the visual prominence of the first portion of the first virtual content is reduced by a first amount, such as 0%, 10%, 20%, 40%, 50%, 70%, or 90%.
[0162] In some embodiments, reducing includes reducing the visual prominence of the second portion of the first virtual content by a second amount (844c) that is greater than the first amount relative to the three-dimensional environment (e.g., 702), such as Figure 7E the reduction in the visual prominence of the portion of the user interface 704 shown. In some embodiments, the visual prominence of the second portion of the first virtual content is reduced by a second amount that is greater than the first amount, such as 30%, 50%, 70%, 90%, or 100%. Thus, in some embodiments, the portion of the first virtual content that is closer to the boundary of the first virtual content is reduced more in terms of visual prominence compared to the portion of the first virtual content that is farther from the boundary of the first virtual content. In some embodiments, the position of the first physical object relative to the first virtual content results in different levels of penetration for the first physical object. For example, if the first physical object is located behind the central region of the first virtual content and triggers a penetration, the computer system reduces the visual prominence of the central region by a first amount (e.g., corresponding to a user-defined setting that defines the amount of penetration), but if the first physical object is located behind the right region of the first virtual content and triggers a penetration, the computer system reduces the visual prominence of the right region by a second amount that is greater than the first amount (e.g., not corresponding to and / or greater than the amount defined by the user-defined setting). The feathered penetration from the edge of the first virtual content reduces the interruption of the main portion (e.g., the central region) of the first virtual content while still allowing the user to interact with the first physical object, which reduces the need for additional input to return to the display of the first virtual content, thereby reducing power usage and extending the battery life of battery-powered devices.
[0163] In some embodiments, in response to detecting a first physical object in a first portion of the physical environment and based on determining that the current value of a first user-defined setting is set to not allow penetration for physical objects that meet one or more first criteria (846a), and based on determining that the first physical object (e.g., 708) at the first location meets one or more second criteria (where one or more of the second criteria are security-related), the visual prominence of a first virtual content (e.g., 704) is reduced with respect to a three-dimensional environment (e.g., 702) (846b). For example, if the current value of the first user-defined setting is set to not allow penetration for physical objects that meet one or more first criteria, and if the person 708 meets one or more second criteria that are security-related, the visual prominence of the user interface 704 is reduced with respect to the three-dimensional environment, such as Figures 7C to 7E shown, even if the person 708 would not otherwise cause penetration. In some embodiments, the computer system defines one or more second criteria that are security-related, such as criteria that are met if conditions in the user's physical environment pose a potential security risk to the user in circumstances where the user cannot see them. For example, if a person or animal is running towards the user of the computer system, an alarm disappears in the user's physical environment, and / or the first physical object is very close to the user of the computer system (e.g., within 0.1 meter, 0.3 meter, 0.5 meter, 1 meter, 3 meters, 5 meters, or 10 meters), then optionally one or more second criteria are met. In such a case, regardless of whether the current value of the first user-defined setting is set to prevent penetration and / or whether one or more first criteria are met, the computer system will optionally reduce the visual prominence of the first virtual content with respect to the three-dimensional environment to allow the user to see potential unsafe conditions in the user's physical environment. Allowing penetration for security reasons protects the user of the computer system from potential security-related issues and protects the device from potential hazards.
[0164] In some embodiments, when a first physical object (e.g., 708) is in a first portion of the physical environment, the computer system (e.g., 101) generates an audio output associated with the three-dimensional environment (e.g., Figure 7A(848a)Audio effects in (e.g., the audio is associated with the first virtual content and / or environmental effects displayed by the computer system, such as audio effects corresponding to a virtual environment such as an ocean environment). In some embodiments, in response to detecting a first physical object (e.g., 708) (848b) in a first portion of the physical environment, the magnitude of the audio output associated with the three-dimensional environment (e.g., 702) is reduced (848c) based on determining that the first physical object (e.g., 708) at the first location meets one or more first criteria and the current value of the first user-defined setting is set to allow penetration of physical objects that meet one or more first criteria (e.g., Figure 7C (848d)Reduction of the audio effects in). In some embodiments, penetration includes reducing the magnitude of the audio output associated with the three-dimensional environment, such as reducing the magnitude by 1%, 10%, 20%, 40%, 60%, 80%, or 100%. In some embodiments, this magnitude is determined by the user or by the computer system. In some embodiments, the magnitude is based on a reduction in the visual prominence of the first virtual content. For example, if the visual prominence of the first virtual content is reduced by 40%, the magnitude will also be reduced by 40% or another amount equal to the amount of reduction in the visual prominence of the first virtual content. Alternatively, the magnitude is independent of the reduction in the visual prominence of the first virtual content. For example, if the visual prominence of the first virtual content is reduced by 40%, the magnitude is optionally reduced by 10% or another amount different from the amount of reduction in the visual prominence of the first virtual content. Reducing the magnitude of the audio output associated with the three-dimensional environment limits distractions caused by the audio output associated with the three-dimensional environment, which reduces the need for additional input to reduce the magnitude of the audio output associated with the three-dimensional environment, thereby reducing power usage and extending the battery life of battery-powered devices.
[0165] (850)In some embodiments, the visual prominence of the first virtual content (e.g., user interface 704) is reduced relative to the three-dimensional environment (e.g., 702). In some embodiments, reducing includes one or more of the following: dimming, blurring, or stopping the display of one or more portions of the first virtual content (e.g., user interface 704). For example, in some embodiments, as Figure 7CAs shown, with respect to the three-dimensional environment 702, the visual prominence of the first virtual content is reduced by dimming, blurring, or stopping the display of one or more portions of the user interface 704. In some embodiments, the first virtual content is optionally dimmed in the three-dimensional environment to reveal a portion of the first physical object. For example, the first virtual content is optionally dimmed to a less distinct shine. Additionally or alternatively, the first virtual content is optionally blurred in the three-dimensional environment to reveal a portion of the first physical object. For example, the first virtual content is optionally displayed less prominently or becomes less distinct. Additionally or alternatively, the first virtual content is optionally hidden in the three-dimensional environment to reveal a portion of the first physical object. Performing the penetration of the first virtual content in various ways ensures that the effectiveness of the penetration is high, which reduces the need for additional input to increase the penetration effectiveness, thereby reducing power usage and extending the battery life of battery-powered devices.
[0166] In some embodiments, one or more of the first criteria include criteria (852) that are not met when the first virtual content is system virtual content. For example, even if the person 710 meets the criteria, the user interface 706 is at Figure 7CNor is it penetrated. In some embodiments, the first virtual content is system virtual content, such as a user interface that displays or presents one or more controls that enable a user to control one or more functionalities of a computer system (e.g., system volume and / or system Wi-Fi). For example, the system virtual content optionally includes play buttons, rewind buttons, fast forward buttons, stop buttons, end buttons, etc. to control the playback of corresponding content (e.g., video or audio). In some embodiments, the computer system does not penetrate the system virtual content, even if the first physical object has characteristics that would otherwise cause the first physical object to penetrate non-system virtual content. In some embodiments, the system virtual content is a system virtual environment (e.g., a simulated three-dimensional environment) that is capable of being displayed in a three-dimensional environment in response to input provided by a user to display such a simulated environment. In some embodiments, the simulated environment includes a scene that at least partially obscures at least a portion of the three-dimensional environment such that it appears as if the user is located within the scene (e.g., and optionally no longer within the three-dimensional environment). In some embodiments, the simulated environment is an atmospheric transformation that modifies one or more visual characteristics of the three-dimensional environment such that it appears as if the three-dimensional environment is located at a different time, place, and / or condition (e.g., morning light instead of afternoon light, sunny instead of cloudy, etc.). In some embodiments, the three-dimensional environment is a mixed environment that includes a simulated environment and a representation of a physical environment. In some embodiments, the boundary between the simulated environment and the representation of the physical environment is based on distance. For example, the portion of the three-dimensional environment that includes the simulated environment is optionally further from the user's viewing point compared to the portion of the three-dimensional environment that includes the representation of the physical environment. Not penetrating the system virtual content ensures an uninterrupted ability to interact with the system virtual content, which reduces the need for additional input to return to the display of the system virtual content, thereby reducing power usage and extending the battery life of battery-powered devices.
[0167] It should be understood that the specific order in which the operations in method 800 are described is merely exemplary and is not intended to indicate that the described order is the only order in which these operations can be performed. Those of ordinary skill in the art will envision various ways to reorder the operations described herein.
[0168] Figures 9A to 9E Illustrates an example of a computer system that displays an indication of a physical object associated with virtual content in a three-dimensional environment according to some embodiments.
[0169] Figure 9A Illustrates a three-dimensional environment 902 visible via a display generation component of computer system 101 (e.g., Figure 1 display generation component 120), the three-dimensional environment 902 being visible from the user's viewing point 926 illustrated in a top view (e.g., which faces the back wall of the physical environment in which computer system 101 is located). As referenced aboveFigures 1 to 6 As described, computer system 101 optionally includes a display generation component (e.g., a touch screen, an HMD or other wearable computing device, a display, a projector, etc.) and a plurality of image sensors (e.g., Figure 3 image sensor 314). The image sensors optionally include one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that computer system 101 can use to capture one or more images of a user or a part of the user (e.g., one or more hands of the user) when the user interacts with computer system 101. In some embodiments, the user interfaces illustrated and described below may also be implemented on a head-mounted display that includes a display generation component for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting movement of the physical environment and / or the user's hand (e.g., external sensors facing out from the user) and / or sensors for detecting the user's gaze (e.g., internal sensors facing in towards the user's face).
[0170] As Figure 9A shown, computer system 101 captures one or more images of the physical environment (e.g., operating environment 100) around computer system 101, including one or more objects in the physical environment around computer system 101. In some embodiments, computer system 101 displays a representation of the physical environment in a three-dimensional environment 902, and / or the physical environment is visible in the three-dimensional environment 902 via display generation component 120. For example, the three-dimensional environment 902 visible via display generation component 120 includes a representation of the physical floor of the room in which computer system 101 is located and the back and side walls. For example, and as Figure 9A shown, the physical environment around computer system 101 includes a couch 924 (shown in a top view), which is visible from Figure 9A the user's viewpoint 926 in.
[0171] In Figure 9A the three-dimensional environment 902 includes user interfaces 904 and 906. In Figure 9A the user interfaces 904 and 906 are two-dimensional objects. It should be understood that the examples of the present disclosure optionally apply equally to three-dimensional objects. User interfaces 904 and 906 are optionally one or more of the following: a user interface of an application (e.g., a messaging user interface and / or a content browsing user interface), a three-dimensional object (e.g., a virtual clock, a virtual ball, and / or a virtual car), or any other element displayed by computer system 101 that is not included in the physical environment of computer system 101. In Figures 9A to 9EIn [the figure], user interface 904 is a content playback user interface (e.g., the user interface through which computer system 101 displays video and / or presents audio content), and user interface 906 is also a content playback user interface. Reference method 1000 provides one or more details of user interfaces 904 and 906.
[0172] In Figure 9A [the figure], computer system 101 is also presenting visual environmental effects to three-dimensional environment 902, such as dimming, desaturating, or otherwise reducing the visual prominence of one or more portions of the physical environment visible via display generation component 120. In some embodiments, computer system 101 additionally or alternatively presents audio environmental effects that at least partially obscure one or more portions of the physical audio that would otherwise be audible from the physical environment. Additional details of the visual and / or audio environmental effects are described in reference methods 800 and / or 1000.
[0173] In some embodiments, computer system 101 presents three-dimensional environment 902 in a manner that occludes or otherwise reduces the visibility of one or more portions of the physical environment surrounding computer system 101. For example, Figure 9A in [the figure], user interface 904 is occluding and / or obscuring a portion of the back wall and floor in the physical environment, and Figure 9A in [the figure], user interface 906 is occluding and / or obscuring a portion of the back wall and floor in the physical environment. As another example, in Figure 9A [the figure], the audio environmental effects and / or visual environmental effects presented by computer system 101 are also occluding and / or obscuring a portion of the couch 924 and / or audio components in the physical environment.
[0174] In some embodiments, and as discussed in more detail below, computer system 101 reduces the visual prominence of one or more virtual portions of three-dimensional environment 902 based on one or more characteristics of objects in the physical environment to allow those physical objects to become more visible in three-dimensional environment 902. In some embodiments, and as discussed in more detail below, computer system 101 presents an indication in three-dimensional environment 902 that represents a corresponding object in the physical environment.
[0175] In Figure 9A [the figure], person 908 and person 910 (two people other than the user of computer system 101) are present within the physical environment surrounding computer system 101. As Figure 9A shown, person 908 and person 910 are located outside the field of view of three-dimensional environment 902 visible via display generation component 120 and are not visible in three-dimensional environment 902.
[0176] FromFigures 9A to 9B , Persons 908 and 910 have moved within the physical environment around computer system 101 and are now partially located behind user interfaces 904 and 906 respectively. As Figure 9B shown, user interface 904 at least partially obscures a portion of Person 908, and computer system 101 presents an indication 909 of the obscured portion of Person 908 in user interface 904 as indication 909 (e.g., as a shadow overlaying user interface 904). Additionally, and as Figure 9B shown, user interface 906 at least partially obscures a portion of Person 910, and computer system 101 presents an indication 911 of the obscured portion of Person 910 in user interface 906 as indication 911 (e.g., as a shadow overlaying user interface 906). This obscuring and subsequent presentation of indications 909, 911 rather than penetrating user interfaces 904 and / or 906 is optionally due to Persons 908 and 910 not exhibiting one or more characteristics corresponding to computer system 101 designating Persons 908 and 910 as participating in significant social interactions with the user of computer system 101, such as described in reference method 800 and as will be described in more detail later. As Figure 9B shown, the sizes of indications 909 and 911 respectively correspond to the sizes of representative portions of Persons 908 and 910. Additionally, as Figure 9B shown, the shapes of indications 909 and 911 respectively correspond to the sizes of representative portions of Persons 908 and 910. In some embodiments, indications such as indications 909 and 911 have a shape and / or size different from that of the representative object (e.g., oval, rectangle, and / or square). However, it should be noted that typically the indications optionally correspond to the shape and / or size of the representative object. In some embodiments, displaying representative indications (e.g., 909 and 911) of the corresponding physical objects allows the user to see and / or determine where the physical objects are within the physical environment without interrupting and / or distracting the user from the display of virtual content (e.g., 904 and 906). Additionally, and because Persons 908 and 910 are not designated as participating in significant social interactions, and as Figure 9B shown, the audio environmental effects and / or visual environmental effects of three-dimensional environment 902 are Figure 9B the same in Figure 9A as in Figure 9B shown, the unobscured portions of Person 908 (e.g., the feet of Person 908) and the unobscured portions of Person 910 (e.g., the feet or Person 910) are visible in three-dimensional environment 902 to which visual environmental effects (e.g., dimming) are applied.
[0177] In some embodiments, an indication representing an object within a physical environment brightens or dims based on the movement of the represented object. For example, when an object moves faster than it previously moved, the indication for that object optionally dims, and when the object moves slower than it previously moved, the indication for that object optionally brightens. From Figures 9B to 9C , person 908 moves faster within the physical environment around computer system 101 and behind user interface 904, and person 910 moves slower within the physical environment around computer system 101 and behind user interface 906. Thus, as Figure 9C shown, the indication 909 for person 910 dims to indicate the faster movement of person 908, and the indication 911 for person 910 brightens to indicate the slower movement of person 908.
[0178] In some embodiments, if an object stops moving, the indication representing the object within the physical environment ceases to be displayed. From Figures 9C to 9D , person 908 moves more slowly within the physical environment, and person 910 stops moving. Thus, and as Figure 9D shown, the indication 909 for person 910 brightens to indicate the slower movement, while the indication 911 for person 908 has disappeared (e.g., is no longer displayed by computer system 101) because person 910 has stopped moving.
[0179] In some embodiments, and as described above, if computer system 101 designates the actions of an object within the physical environment as exhibiting characteristics of significant social interaction, the corresponding object will penetrate the three-dimensional environment in one or more ways described with reference to method 800 and / or 1000 and described below. Additionally, if an object is located in front of a user interface (e.g., user interface 904 or user interface 906) within the three-dimensional environment 902, computer system 101 optionally changes one or more visual characteristics of the portion of the user interface obscured by the object (e.g., blurs and / or fades that portion of the object).
[0180] From Figures 9D to 9E , person 908 exhibits one or more characteristics corresponding to computer system 101 designating person 908 as participating in significant social interaction with a user of computer system 101, and person 910 moves in front of user interface 906. Reference methods 800 and / or 1000 provide details regarding the various criteria used by computer system 101 to designate such significant social interaction. For example, in Figure 9E , person 908 has turned and / or is talking to the user of computer system 101. Thus, person 908 is eligible to penetrate user interface 904, as Figure 9E shown. Thus, and as Figure 9EAs shown, person 908 penetrates user interface 904. As previously mentioned, person 908 penetrating user interface 904 and / or the three-dimensional environment 902 optionally has one or more of the penetration characteristics described with reference to method 800. Thus, as Figures 9D to 9E shown, computer system 101 optionally transitions from displaying an indication of person 908 associated with user interface 904 to penetrating user interface 904 based on determining that person 908 meets the penetration criteria.
[0181] In addition, and as Figure 9E shown, the head 912 of person 910 is obscuring user interface 906, while the body of person 910 is below user interface 906 within the three-dimensional environment 902 (e.g., not obscuring user interface 906). In some embodiments, and as Figure 9E shown, in response to such obscuring, computer system 101 modifies one or more visual characteristics of the head of person 910 to present a representation of head 912 of person 910 in a blurred, faded, less saturated, and / or more transparent manner and / or in other ways having reduced visual prominence relative to the three-dimensional environment 902. Thus, the obscuring of user interface 906 is optionally reduced, and portions of user interface 906 otherwise obscured by the representation of the head 912 of person 910 are more visible within the three-dimensional environment 902. Additionally, as Figure 9E shown, computer system 101 does not modify the visual characteristics of the body of person 910, and thus the body of person 910 is visible within the three-dimensional environment 902 to which visual environment effects (e.g., dimming) are applied. It should be noted that portions of user interface 906 not obscured by the head of person 910 maintain their visual prominence such that the user can still view those portions of user interface 906. Additionally, and as Figure 9E shown, the visual appearance of the region and / or volume external to person 910 (e.g., the three-dimensional environment 902 including the unobscured portions of user interface 906) does not change.
[0182] Figures 10A to 10F is a flowchart illustrating method 1000 for displaying an indication of a physical object in a three-dimensional environment according to some embodiments. In some embodiments, method 1000 is performed at a computer system (e.g., Figure 1 computer system 101 in, such as a tablet device, a smart phone, a wearable computer, or a head-mounted device), the computer system including a display generation component (e.g., Figure 1 , Figure 3 and Figure 4a display generation component 120) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 1000 is managed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., Figure 1 the controller 110) in). Some operations in method 1000 are optionally combined, and / or the order of some operations is optionally changed.
[0183] In some embodiments, method 1000 is executed at a computer system (e.g., computer system 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., 314). In some embodiments, the computer system has one or more features of the computer system in method 800. In some embodiments, the display generation component has one or more features of the display generation component of method 800. In some embodiments, the one or more input devices have one or more of the features of the one or more input devices of method 800.
[0184] In some embodiments, when a three-dimensional environment (e.g., 902) is visible via a display generation component (e.g., 120) (e.g., the three-dimensional environment has one or more of the characteristics of the three-dimensional environment of method 800), and the three-dimensional environment (e.g., 902) includes first virtual content (e.g., the first virtual content has one or more of the characteristics of the first virtual content of method 800) that is obscuring a first portion of the physical environment (1002a), the computer system (e.g., 101) detects a first physical object (e.g., 908) located in the first portion of the physical environment (e.g., a person other than the user, an animal such as a pet, or any other physical object capable of moving) (1002b). In some embodiments, the first portion of the physical environment is obscured in a manner similar to that described with respect to method 800. In some embodiments, the first physical object is detected via one or more input devices in a manner similar to that described with respect to method 800. In some embodiments, in response to detecting the first physical object (e.g., 908) in the first portion of the physical environment, and based on determining that the first physical object (e.g., 908) meets one or more first criteria (e.g., satisfaction of the one or more first criteria corresponds to significant social interaction, and the one or more criteria are discussed in more detail below and / or with reference to method 800), the computer system (e.g., 101) reduces the visual prominence of the first virtual content relative to the three-dimensional environment (e.g., 902) to reveal a portion of the first physical object (e.g., 908) (1002c), e.g., reduces the visual prominence of the user interface 904 relative to the three-dimensional environment 902 to reveal a portion of the person 908 (e.g., including increasing the visibility of the first portion of the physical environment through the first virtual content). In some embodiments, the visual prominence of the first virtual content (or a portion of the first virtual content) is reduced in a manner similar to that described with respect to method 800 to reveal a portion of the first physical object. In some embodiments, the opacity, brightness, color saturation, and / or other visual characteristics of the first virtual content are optionally reduced or changed by a first amount to increase the visibility of the first physical object through the first virtual content (e.g., increasing visibility from 0% visible to 20%, 40%, 60%, or 80% visible, or from 5% visible to 20%, 40%, 60%, or 80% visible).
[0185] In some embodiments, based on determining that the first physical object (e.g., 908) does not meet the one or more first criteria, the computer system (e.g., 101) concurrently displays an indication (e.g., 909) of a portion of the first physical object (e.g., 908) with the first virtual content (e.g., 904) in the three-dimensional environment (e.g., 902) without revealing the portion of the first physical object (e.g., 1002d), e.g., in Figure 9BIn this case, the indication 909 of the person 908 is concurrently displayed in the three-dimensional environment 902 with the user interface 904 without revealing portions of the person 908 (e.g., without increasing the visibility of a first portion of the physical environment through the first virtual content). In some embodiments, the computer system displays an indication of the first physical object and / or first virtual content corresponding to its location within the physical environment at a location within the three-dimensional environment. For example, rather than reducing or changing the opacity, brightness, color saturation, and / or other visual characteristics of the first virtual content to increase the visibility of the first physical object through the first virtual content, the computer system optionally displays a shadow of the first physical object that overlays and / or passes through a portion of the first virtual content that is positioned between the user's viewpoint and the portion of the first physical object within the three-dimensional environment. This indication of the first physical object is optionally displayed in a variety of ways, which are discussed in more detail hereinafter, such as with reference to steps 1016 through 1020. In some embodiments, the computer system reduces or changes the opacity, brightness, color saturation, and / or other visual characteristics of the first virtual content by a second amount that is less than a first amount (e.g., the first amount is a 100%, 75%, 50%, 40%, or 30% reduction in visual prominence, and the second amount is a 50%, 40%, 20%, 10%, or 5% reduction in visual prominence) such that the indication of the portion of the first physical object is visible through the first virtual content. In some embodiments, the brightness of the first virtual content (optionally only at locations corresponding to the portion of the first physical object) is reduced (optionally without reducing the opacity) as an indication of the portion of the first physical object. When the penetration criterion is not met, selectively reducing the visual prominence of the virtual content or enabling the display of the indication allows the user to interact with the physical environment when needed while maintaining the display of the virtual content in a manner that still allows the user to otherwise be aware of the physical environment and reduces distraction to the user, which reduces input errors and makes interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0186] In some embodiments, a first physical object (e.g., 908) is a first type of physical object (1004a) (e.g., the first physical object is a living object such as a person or an animal), and a second physical object (e.g., 924) (1004b) located in a first portion of the physical environment is detected via one or more input devices. In some embodiments, in response to detecting the second physical object (e.g., 924) in the first portion of the physical environment (1004c), and based on determining that the second physical object (e.g., 924) is a second type of physical object that is different from the first type of physical object (e.g., the second physical object is an inanimate object such as a couch, a chair, and / or a table), the computer system (e.g., 101) maintains the visual prominence of the first virtual content relative to the three-dimensional environment (e.g., 902) without revealing a portion of the second physical object (e.g., 924) and without displaying an indication of a portion of the second physical object with the first virtual content (1004d). For example, if in Figure 9B a person 908 is a different object of the second type and the computer system does not display the indication 909. In some embodiments, and upon determining that the second physical object is an inanimate object located in the first portion of the physical environment, the visual prominence of the first virtual content is maintained relative to the three-dimensional environment; additionally, optionally, an indication of the second physical object is not displayed. By not allowing penetration or indication of inanimate objects, the user of the computer system can interact with the three-dimensional environment with reduced distraction, which reduces input errors and makes interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0187] In some embodiments, one or more first criteria include criteria (1006) that are satisfied based on the movement of the first physical object (e.g., 908). For example, if a person is moving within the physical environment, the criteria are satisfied and an indication (e.g., 909) is displayed, such as Figure 9BAs shown. In some embodiments, if a person or an animal moves, the action meets the criteria. In some embodiments, before the criteria are met, movement must be detected within a pre-determined time period (e.g., 0.5 seconds, 1 second, 3 seconds, 5 seconds, 10 seconds, 20 seconds, or 60 seconds). The time period is optionally set by the user or determined by the computer system. In some embodiments, if a person or an animal does not move, the criteria are not met. In some embodiments, for the criteria to be met, the movement must be greater than a threshold distance (e.g., 0.05 meters, 0.1 meters, 0.3 meters, 0.5 meters, 1 meter, 3 meters, 5 meters, 10 meters, or 20 meters), have a speed greater than a threshold speed (e.g., 0.1 m / s, 0.3 m / s, 0.5 m / s, 1 m / s, 3 m / s, 5 m / s, or 10 m / s), and / or have an acceleration greater than a threshold acceleration (e.g., 0.01 m / s², 0.05 m / s², 0.1 m / s², 0.3 m / s², 0.5 m / s², 1 m / s², 3 m / s², 5 m / s², or 10 m / s²). Requiring the object movement to meet the criteria reduces the likelihood that an uninteresting object (e.g., a stationary inanimate object) will interrupt the user interaction with the first virtual content, which reduces the input required to redisplay the first virtual content and makes the interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0188] In some embodiments, displaying an indication (e.g., 909) of a portion of a first physical object (e.g., 908) includes dimming at least a portion of a first virtual content (e.g., 904) corresponding to the first physical object (e.g., 908), wherein the indication (e.g., 909) includes at least a portion (1008) of the dimming of the first virtual content (e.g., 904), such as Figure 9B , Figure 9C and Figure 9D the indication 909 shown. In some embodiments, the indication is displayed as a dimmed portion of at least a portion of the first virtual content. The darkness of the indication can optionally vary based on various factors, which will be discussed in more detail below, such as with reference to steps 1016 to 1020. Dimming a portion of the content enables the user to view the first virtual content with limited distraction and still be informed of the physical object in the physical environment, which reduces input errors and makes the interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0189] In some embodiments, the indication (e.g., 909) of a portion of the first physical object (e.g., 908) has a size and / or shape (1010) corresponding to the size and / or shape of the first physical object (e.g., 908), such as Figure 9B , Figure 9C and Figure 9DThe indication 909 shown. In some embodiments, the indication is generally shaped to resemble the first physical object. For example, if the first physical object is a tall person, the indication will generally be shaped like a tall person, and the indication will cause a larger portion of the first virtual content to darken. Alternatively, if the first physical object is a small dog, the indication will generally be shaped like a small dog, and the indication will cause only a small portion of the first virtual content to darken. In some embodiments, the indication has the same shape as the first physical object. In some embodiments, the shape of the indication corresponds to the shape of the first physical object in the same or a similar manner as the shape of a shadow corresponds to the object casting the shadow. In some embodiments, the indication has a shape that is an approximation of the shape of the physical object (e.g., an oval or an irregular blotch the size of a person or a dog). Displaying an indication having the general shape of the object enables the user of the computer system to more easily determine what physical object the indication represents, which reduces input errors and makes interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0190] In some embodiments, an indication (e.g., 909) of a portion of a first physical object (e.g., 908) is displayed at a first location (1012a) relative to a first virtual content (e.g., 904) (e.g., the indication is displayed at a first location within and / or overlaid on the first virtual content, such as in the right hand portion of the first virtual content, because the first physical object is located behind the right hand portion of the first virtual content). In some embodiments, when an indication (e.g., 909) of a first portion of a first physical object (e.g., 908) is displayed at a first location relative to a first virtual content (e.g., 904), a movement of the first physical object (e.g., 904) from a second location to a third location in a first portion of the physical environment is detected via one or more input devices (e.g., 314) (e.g., a movement of the first physical object from behind the right hand portion of the first virtual content to behind the left hand portion of the first virtual content). In some embodiments, in response to (and / or when) detecting the movement of the first physical object (e.g., 908) from the second location to the third location in the first portion of the physical environment, the computer system moves (1012c) the indication (e.g., 909) of the first portion of the first physical object (e.g., 908) from the first location relative to the first virtual content to a fourth location different from the first location relative to the first virtual content (e.g., the indication is displayed at a second location within the first virtual content and / or overlaid on the first virtual content, such as in the left hand portion of the first virtual content, because the first physical object is located behind the left hand portion of the first virtual content). In some embodiments, the fourth location of the indication relative to the first virtual content corresponds to the third location (1012c) of the first physical object in the first portion of the physical environment. For example, as Figure 9B , Figure 9C and Figure 9DAs shown, the indication 909 of the person 908 moves within the three-dimensional environment 902 in a manner that mirrors the movement of the person 908 within the physical environment as shown in the top view. In some embodiments, the indication moves within the three-dimensional environment and / or within the first virtual content in a manner that mirrors the movement of the first physical object within the physical environment. For example, if the first physical object moves from a second position to a third position within the physical environment, the corresponding indication within the three-dimensional environment will likewise move from the corresponding second position within the three-dimensional environment to the corresponding third position within the three-dimensional environment. In some embodiments, the properties of the indication are optionally changed based on the movement of the first physical object, which is discussed in more detail below with reference to steps 1016 to 1018. Moving the indication within the three-dimensional environment based on the movement of the corresponding first physical object within the physical environment gives the user of the computer system feedback on where the first physical object is moving and at what speed, which reduces input errors and makes interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0191] In some embodiments, the computer system (e.g., 101) reduces the visual prominence of the first virtual content relative to the three-dimensional environment (e.g., 902), which includes reducing the visual prominence of the first part (e.g., 904) of the first virtual content, where the first part of the first virtual content has a first size (1014a). In some embodiments, instead of displaying an indication of the first physical object, the visual prominence of the first part of the first virtual content is reduced (e.g., penetrated), and the penetrated part of the first virtual content has the first size. In some embodiments, the indication (e.g., 909) of the part of the first physical object has a second size (1014b) that is smaller than the first size. For example, as Figure 9EAs shown, if person 908 penetrates user interface 904, indication 909 optionally has a smaller size / representation in three-dimensional environment 902 than the penetration area. In some embodiments, and in the case where an indication of the portion of the first physical object is displayed within the three-dimensional environment, the indication has a smaller size than the size of the penetration. For example, if the first physical object meets the penetration criteria and is visible in the three-dimensional environment, the penetrated portion of the first virtual content (e.g., as described in more detail with reference to method 800) is larger than the indication of the first physical object displayed in the case where the first physical object does not meet the penetration criteria but meets the indication criteria. In some embodiments, the (general) shape of the penetrated portion of the first virtual content is the same as the (general) shape of the indication of the first physical object displayed, which optionally corresponds to the shape and / or size of the first physical object, as described above. In some embodiments, the (general) shape of the penetrated portion of the first virtual content is different from the (general) shape of the indication of the first physical object displayed (e.g., is rectangular rather than having a shape that mirrors the shape of the first physical object); instead, the size of the penetrated portion of the first virtual content optionally corresponds to the size of the first physical object, but the shape of the penetrated portion of the first virtual content optionally does not correspond to the shape of the first physical object. Making the penetration size of the first physical object larger than the size of the indication representing the first physical object makes the user more likely to be aware of the penetration than the indication of the physical object, which reduces input errors and makes interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0192] In some embodiments, in response to detecting a first physical object (e.g., 908) in a first portion of a physical environment and based on determining that the first physical object (e.g., 908) does not meet one or more first criteria, a computer system (e.g., 101) displays an indication (e.g., 909) of a portion of the first physical object (e.g., 908) with a first visual prominence (1016a) (e.g., a first brightness, a first opacity, a first saturation, and / or a first contrast level relative to the remainder of the first virtual content). In some embodiments, when displaying the indication (e.g., 909) of the portion of the first physical object (e.g., 908) with the first visual prominence, a decrease in the magnitude of movement (e.g., distance, speed, and / or acceleration) of the first physical object (e.g., 908) is detected (1016b). In some embodiments, in order for the decrease in magnitude to be noticed by the computer system, the movement of the first physical object must be decreased by at least a predetermined magnitude, as described in more detail below. The size is optionally set by a user of the computer system or determined by the computer system. In some embodiments, in response to detecting the decrease in the magnitude of movement of the first physical object (e.g., 908), the visual prominence of the indication of the portion of the first physical object (e.g., 908) is decreased (1016c) (e.g., to a second brightness greater than the first brightness, a second opacity less than the first opacity, a second saturation less than the first saturation, and / or a second contrast level less than the first contrast level relative to the remainder of the first virtual content). For example, as Figure 9C shown, when person 910 moves quickly, indication 909 of person 908 darkens, and when person 908 moves slowly, indication 911 of person 910 brightens. In some embodiments, and in response to detecting that the magnitude of movement of the first physical object has been decreased by a predetermined magnitude (5%, 10%, 20%, 40%, 50%, 70%, or 90%), the visual prominence of the indication of that portion of the first physical object is decreased. For example, if the first physical object slows down, the indication will optionally brighten. Alternatively, if the first physical object speeds up, the indication will optionally darken. In some embodiments, if the movement of the object does not decrease by the predetermined magnitude, the visual prominence of the indication is maintained. Making the indication of the first physical object brighten in response to the decrease in the magnitude of movement of the first physical object gives the user an indication of the speed at which the first physical object is moving, reduces the interruption of the interaction with the first virtual content, and reduces the interruption caused by stationary objects that the user is less likely to be interested in, which reduces the input required to display the first virtual content and makes the interaction with the device more efficient, thereby reducing power usage and extending the battery life of battery-powered devices.
[0193] In some embodiments, reducing the visual prominence of an indication (e.g., 911) of a portion of a first physical object (e.g., 910) includes ceasing the display (1018) of the indication (e.g., 911) of the portion of the first physical object (e.g., 910), e.g., if the person 910 stops moving, as Figure 9D shown, the indication 911 of the person 910 will optionally cease to be displayed (e.g., if the first physical object does not meet the penetration criteria, optionally not penetrate the first virtual content). In some embodiments, and in response to detecting a 100% decrease in the magnitude of movement of the first physical object (e.g., the first physical object stops moving), the display of the indication of the portion of the first physical object is ceased. In some embodiments, if the display of the indication of the portion of the first physical object ceases, the first virtual content is displayed in the same manner (e.g., the same visual prominence and / or other visual characteristics) as if there were no physical object in the first portion of the physical environment. Ceasing the display of the indication of the first physical object in response to the cessation of movement of the first physical object gives the user an indication that the first physical object has stopped moving and that the interruption of interaction with the first virtual content is further reduced, which reduces the input required to display the first virtual content and makes interaction with the device more efficient, thereby reducing power usage and extending the battery life of a battery-powered device.
[0194] In some embodiments, reducing the visual prominence of a first virtual content (e.g., 904) relative to a three-dimensional environment includes ceasing the display (1020) of at least a portion of the first virtual content (e.g., 904), e.g., Figure 9E a reduction in the visual prominence of the user interface 904 in
[0195] In some embodiments, when it is determined that a first physical object (e.g., 908) does not meet one or more first criteria, an indication (e.g., 909) of a portion of the first physical object (e.g., 908) is concurrently displayed with first virtual content (e.g., 904) in a three-dimensional environment (e.g., 902) without revealing the portion of the first physical object (e.g., 908), a determination is made that the first physical object (e.g., 908) meets one or more first criteria (1022a). In some embodiments, in response to determining that the first physical object (e.g., 908) meets one or more first criteria, the computer system stops displaying the indication (e.g., 909) of the portion of the first physical object (e.g., 908) (1022b), e.g., person 908 penetrates Figure 9E user interface 904 in and the computer system 101 no longer displays the indication 909 (and optionally also penetrates the first virtual content, as described above and in method 800). For example, the first physical object is optionally a person walking around the physical environment but not directly talking to the user of the computer system, and thus, an indication of the person is displayed within the three-dimensional environment (e.g., on or within the first virtual content), but the first virtual content is not penetrated. However, if the person optionally turns and starts directly talking to the user of the computer system, one or more first criteria are met, and the indication stops being displayed and / or the person penetrates the first virtual content. Alternatively, if one or more first criteria remain unmet, the display of the indication is optionally maintained. Stopping the display of the indication if the penetration criteria are met reduces the resources required to unnecessarily display the indication in the environment, thereby reducing power usage and extending the battery life of battery-powered devices.
[0196] In some embodiments, when the first physical object (e.g., 908) is in a first portion of the physical environment and when an indication (e.g., 909) of a portion of the first physical object (e.g., 908) is not being displayed (e.g., because the first physical object is not moving), it is detected via one or more input devices (e.g., 314) that the first physical object (e.g., 908) meets one or more first criteria (1024a). In some embodiments, in response to detecting that the first physical object (e.g., 908) meets one or more first criteria, the computer system (e.g., 101) reduces the visual prominence of the first virtual content (e.g., 904) relative to the three-dimensional environment (e.g., 902) to reveal the portion of the first physical object (e.g., 908) (1024b). For example, even if no indication (e.g., 909) is being displayed for person 908 and person 908 meets one or more first criteria, person 908 penetrates, as Figure 9EAs shown. In some embodiments, even if no indication is displayed for the first physical object, but it is detected that the first physical object meets one or more first criteria, the visual prominence of the first virtual content is reduced relative to the three-dimensional environment to reveal that portion of the first physical object. For example, if the first physical object is optionally a person sitting on a couch (and thus no indication representing the person is displayed) and the person turns towards the user of the computer system and starts talking directly to the user of the computer system, then the person will optionally penetrate the first virtual content. Even if the indication is not displayed, if the first physical object meets the penetration criteria, revealing that portion of the first physical object also enables the user to view and interact with the first physical object in a timely manner, which reduces the need for additional input to penetrate the first physical object, thereby reducing power consumption and extending the battery life of battery-powered devices.
[0197] In some embodiments, the first physical object (e.g., 908) is a person other than the user of the computer system (e.g., 101) (e.g., the user's spouse, partner, child, friend, or just another person in the environment), and one or more first criteria include criteria that are met when the person's gaze is directed towards the user (e.g., of the computer system), the person (e.g., 908) is oriented towards the user, the person (e.g., 908) is talking to the user, the person (e.g., 908) is a known person designated by the user, or the person (e.g., 908) is within a threshold distance of the user (1026). For example, if person 908 gazes at the user of computer system 101, is oriented towards the user of computer system 101, is talking to the user of computer system 101, is a known person to the user of computer system 101, and / or is within a threshold distance of the user of computer system 101, then person 908 will optionally penetrate, as Figure 9E shown. Details of such criteria are optionally as described with reference to method 800. Allowing penetration based on the gaze of a person directed towards the user of the computer system, the orientation of the person relative to the user of the computer system, the person talking to the user of the computer system, the computer system's recognition of the person, and / or the position of the person relative to the user of the computer system reduces the amount of input required for the user to interact with the person, thereby reducing power consumption and extending the battery life of battery-powered devices.
[0198] In some embodiments, when a first virtual content (e.g., 904) is displayed with reduced visual prominence and it is determined that a first physical object (e.g., 908) meets one or more first criteria, it is detected that the first physical object (e.g., 908) no longer meets the one or more first criteria (1028a). In some embodiments, the first physical object is optionally a person who meets one or more first criteria by directly communicating with a user of the computer system to achieve penetration. In some embodiments, the person stops directly communicating with the user of the computer system and thus no longer meets the one or more first criteria. In some embodiments, in response to detecting that the first physical object (e.g., 908) no longer meets the one or more first criteria (1028b), the visual prominence of the first virtual content (e.g., 904) is increased relative to the three-dimensional environment (1028c). For example, if person 908 is penetrated as shown in Figure 9E and subsequently no longer meets the one or more first criteria, the visual prominence of the user interface 904 is increased, as shown in Figure 9A . In some embodiments, and in response to the person no longer meeting the one or more first criteria, the computer system reduces the penetration of the person such that the user of the computer system can better view the first virtual content. For example, the visual prominence of the first virtual content is increased back to its visual prominence prior to penetration. In some embodiments, an indication (e.g., 909) of a portion of the first physical object (e.g., 908) is concurrently displayed with the first virtual content (e.g., 904) in the three-dimensional environment (e.g., 902) without revealing the portion of the first physical object (e.g., 908) (1028d). For example, an indication 909 of person 908 is concurrently displayed with the user interface 904 in the three-dimensional environment 902 without revealing person 908, as shown in Figure 9B . In some embodiments, an indication is displayed within the three-dimensional environment to indicate a person who no longer meets the penetration criteria. If the physical object no longer meets the penetration criteria, stopping the penetration and displaying the indication reduces distraction for the user of the computer system, which reduces the need for additional input to return to the display of the indication, thereby reducing power usage and extending the battery life of a battery-powered device.
[0199] In some embodiments, in response to detecting a first physical object (e.g., 910) (1030a) in a first portion of a physical environment, and based on determining that the first physical object (e.g., 910) is in front of a first virtual content (e.g., 906) and a portion of the physical object (e.g., 910) is occluding at least a portion of the first virtual content (e.g., in some embodiments, optionally from the perspective of the user's viewpoint, the first physical object is in front of the first virtual content. For example, the first physical object is optionally a person walking in front of at least a portion of the first virtual content.), modify the visual appearance (1030b) of the portion of the first physical object (e.g., 910) that is occluding at least a portion of the first virtual content (e.g., 906), such as Figure 9E the modification of the head 912 of the person 910 in Figure 9E . For example, the computer system performs one or more of the following to change the appearance of the portion of the first physical object that is occluding the first virtual content from the user's viewpoint: 1) increase the transparency of the portion of the first physical object such that the first virtual content is at least partially visible through the portion of the first physical object; 2) increase the blurriness of the portion of the first physical object; 3) decrease the color saturation of the portion of the first physical object; 4) decrease the brightness of the portion of the first physical object. Modifying the appearance of the portion of the object that is occluding the virtual content indicates to the user that the virtual content is being occluded and reduces the impact of such occlusion on the user, which reduces the need for additional input to reduce occlusion, thereby reducing power usage and extending the battery life of battery-powered devices.
[0200] It should be understood that the particular order in which the operations in method 1000 are described is merely exemplary and is not intended to indicate that the described order is the only order in which these operations can be performed. Those of ordinary skill in the art will envision various ways to reorder the operations described herein. In some embodiments, aspects / operations of methods 800 and 1000 can be interchanged, substituted, and / or added between these methods. For example, the three-dimensional environment of methods 800 and / or 1000, the indication of objects in methods 800 and / or 1000, the environmental effects of methods 800 and / or 1000, the penetration criteria of methods 800 and / or 1000, and / or the characteristics of physical objects in methods 800 and / or 1000 are optionally interchanged, substituted, and / or added between these methods. For the sake of brevity, these details are not repeated here.
[0201] For purposes of explanation, the foregoing description has been presented by reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, to thereby enable others skilled in the art to best utilize the invention in various embodiments with various modifications as are suited to the particular use contemplated.
[0202] As described above, one aspect of the present technology is to collect and use data obtained from various sources to improve a user's XR experience. The present disclosure contemplates that, in some instances, the collected data may include personal information data that uniquely identifies or can be used to contact or locate a specific person. Such personal information data can include demographic data, location-based data, phone numbers, email addresses, Twitter IDs, home addresses, data or records relating to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.
[0203] The present disclosure recognizes that the use of such personal information data in the technology of the present invention can be used to benefit the user. For example, personal information data can be used to improve a user's XR experience. Additionally, the present disclosure contemplates other uses of personal information data that are beneficial to the user. For example, health and fitness data can be used to provide insights into a user's overall health condition or can be used as positive feedback to individuals who use technology to pursue health goals.
[0204] The present disclosure anticipates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with sound privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and privacy practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. Such policies should be accessible to users and should be updated as the collection and / or use of data changes. Personal information from users should be collected for legitimate and reasonable entity purposes and should not be shared or sold outside of those legitimate purposes. Additionally, such collection / sharing should occur after receiving informed consent from the user. Further, such entities should consider taking any necessary steps to protect and secure access to such personal information data and to ensure that other entities with access to personal information data comply with the privacy policies and procedures of the other entities. Additionally, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and privacy practices. Moreover, the policies and practices should be adapted to the specific types of personal information data being collected and / or accessed and to the applicable laws and standards, including considerations of particular jurisdictions. For example, in the United States, the collection or acquisition of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while health data in other countries may be subject to other regulations and policies and should be handled accordingly. Thus, different privacy practices should be asserted for different types of personal data in each country.
[0205] Notwithstanding the foregoing, the present disclosure also anticipates embodiments where users selectively block the use or access of personal information data. That is, the present disclosure anticipates that hardware elements and / or software elements may be provided to prevent or block access to such personal information data. For example, in the context of XR experiences, the inventive technology can be configured to allow users to select "opt-in" or "opt-out" of participating in the collection of personal information data during or at any time after the registration service. In addition to providing "opt-in" and "opt-out" options, the present disclosure also anticipates providing notifications related to the access or use of personal information. For example, the user may be notified when downloading an application that the user's personal information data will be accessed and then reminded again just before the personal information data is accessed by the application.
[0206] In addition, it is the intention of the present disclosure that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once the data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. Additionally, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated, when appropriate, by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of the data stored (e.g., collecting location data at the city level rather than at the address level), controlling how the data is stored (e.g., aggregating data among users), and / or other methods.
[0207] Accordingly, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that the various embodiments may also be implemented without access to such personal information data. That is, the various embodiments of the inventive technology will not fail to operate properly due to the lack of all or a portion of such personal information data. For example, an XR experience can be generated by inferring preferences based on non-personal information data or an absolute minimum amount of personal information (such as content requested by a device associated with the user, other non-personal information available to the service, or publicly available information).
Claims
1. A method, comprising: At a computer system in communication with a display generation component and one or more input devices: When a three-dimensional environment is visible via the display generation component, detecting, via the one or more input devices, a first physical object at a first location in a first portion of a physical environment, wherein the three-dimensional environment includes first virtual content that is obscuring the first portion of the physical environment; and In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first location meets one or more first criteria and a current value of a first user-defined setting is set to allow penetration for physical objects that meet the one or more first criteria, reducing a visual prominence of the first virtual content relative to the three-dimensional environment; And Based on determining that the first physical object at the first location meets the one or more first criteria while the current value of the first user-defined setting is set to prevent penetration for physical objects that meet the one or more first criteria, maintaining the visual prominence of the first virtual content relative to the three-dimensional environment.
2. The method according to claim 1, the method further comprising: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first location does not meet the one or more first criteria, maintaining the visual prominence of the first virtual content relative to the three-dimensional environment.
3. The method according to claim 1, wherein when the current value of the first user-defined setting corresponds to a do-not-disturb mode of the computer system, the current value of the first user-defined setting is set to prevent penetration for physical objects that meet the one or more first criteria, and during the do-not-disturb mode, one or more notifications are suppressed in response to detecting the generation of one or more notification events.
4. The method according to claim 1, the method further comprising: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the current value of the first user-defined setting corresponds to a dynamic setting for the first user-defined setting and determining that the one or more first criteria are met, reducing the visual prominence of the first virtual content relative to the three-dimensional environment, wherein satisfaction of the one or more first criteria is based on one or more characteristics of the first physical object; And Based on determining that the current value of the first user-defined setting corresponds to a dynamic setting for the first user-defined setting and determining that the one or more first criteria are not met, maintaining the visual prominence of the first virtual content relative to the three-dimensional environment.
5. The method according to claim 4, wherein the first physical object is a person other than a user of the computer system, and the one or more first criteria include criteria that are met when the person's gaze is directed at the user of the computer system.
6. The method according to claim 4, wherein the first physical object is a person other than the user of the computer system, and the one or more first criteria include criteria that are satisfied when the person is oriented towards the user of the computer system.
7. The method according to claim 4, wherein the first physical object is a person other than the user of the computer system, and the one or more first criteria include criteria that are satisfied when the person is talking to the user of the computer system.
8. The method according to claim 4, wherein the first physical object is a person other than the user of the computer system, and the one or more criteria include criteria that are satisfied when the person is a known person as specified by the user of the computer system.
9. The method according to claim 4, wherein the first physical object is a person other than the user of the computer system, and the one or more first criteria include criteria that are satisfied when the person is within a threshold distance of the user of the computer system.
10. The method according to claim 1, wherein the three-dimensional environment includes the first virtual content and the three-dimensional environment is displayed with environmental effects, and the method further includes: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first position satisfies the one or more first criteria and the current value of the first user-defined setting is set to allow penetration of physical objects that satisfy the one or more first criteria, reducing the visual prominence of the environmental effects in the three-dimensional environment.
11. The method according to claim 10, wherein the environmental effects are displayed in a first portion of the three-dimensional environment, the first portion being within a threshold distance of the projection of the first physical object relative to the viewpoint of the user of the computer system, and the environmental effects are displayed in a second portion of the three-dimensional environment, the second portion being outside the threshold distance of the projection of the first physical object relative to the viewpoint of the user, and the method further includes: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first position satisfies the one or more first criteria and the current value of the first user-defined setting is set to allow penetration of physical objects that satisfy the one or more first criteria, reducing the visual prominence of the first virtual content and the visual prominence of the environmental effects in the first portion of the three-dimensional environment, while maintaining the visual prominence of the environmental effects in the second portion of the three-dimensional environment.
12. The method according to claim 1, wherein a first portion of the first virtual content is displayed in a first portion of the three-dimensional environment, the first portion being within a threshold distance of the projection of the first physical object relative to the viewpoint of the user of the computer system, and a second portion of the first virtual content is displayed in a second portion of the three-dimensional environment, the second portion being outside the threshold distance of the projection of the first physical object relative to the viewpoint of the user, and the method further comprises: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first location meets the one or more first criteria and the current value of the first user-defined setting is set to allow penetration for physical objects meeting the one or more first criteria, reducing the visual prominence of the first portion of the first virtual content in the first portion of the three-dimensional environment while maintaining the visual prominence of the second portion of the first virtual content in the second portion of the three-dimensional environment.
13. The method according to claim 1, wherein the three-dimensional environment further comprises second virtual content, wherein the first physical object is located between the first virtual content and the second virtual content relative to the viewpoint of the user of the computer system, and the second virtual content is farther from the viewpoint of the user than the first virtual content, and the method further comprises: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first location meets the one or more first criteria and the current value of the first user-defined setting is set to allow penetration for physical objects meeting the one or more first criteria, reducing the visual prominence of the first virtual content while maintaining the visual prominence of the second virtual content relative to the three-dimensional environment.
14. The method according to claim 1, wherein the three-dimensional environment further comprises second virtual content, wherein the first physical object is located between the first virtual content and the second virtual content relative to the viewpoint of the user of the computer system, and the second virtual content is farther from the viewpoint of the user than the first virtual content, and the method further comprises: In response to detecting the first physical object in the first portion of the physical environment: Based on determining that the first physical object at the first location meets the one or more first criteria and the current value of the first user-defined setting is set to allow penetration for physical objects meeting the one or more first criteria, reducing the visual prominence of the first virtual content and reducing the visual prominence of the second virtual content relative to the three-dimensional environment.
15. The method according to claim 1, the method further comprises: In response to detecting the first physical object in the first part of the physical environment: Based on determining that the current value of the first user-defined setting is set to a first value, and based on determining that one or more second criteria are met, reducing the visual prominence of the first virtual content by a first amount with respect to the three-dimensional environment; And Based on determining that the current value of the first user-defined setting is set to a second value different from the first value, and Based on determining that one or more third criteria are met, reducing the visual prominence of the first virtual content by a second amount with respect to the three-dimensional environment.
16. The method according to claim 15, wherein the one or more second criteria are different from the one or more third criteria.
17. The method according to claim 15, wherein the second amount is different from the first amount.
18. The method according to claim 15, the method further comprising: Detecting, via the one or more input devices, user input corresponding to a request to set the current value of the first user-defined setting to the first value, wherein the user input is detected when the current value of the first user-defined setting corresponds to a dynamic setting for the first user-defined setting; In response to detecting the user input, setting the current value of the first user-defined setting to the first value; and After setting the current value of the first user-defined setting to the first value, automatically adjusting the current value of the first user-defined setting to correspond to the dynamic setting based on determining that a period of time has elapsed since receiving the user input.
19. The method according to claim 18, wherein the user input corresponds to a request to switch between a first mode and a second mode, in the first mode, the current value of the first user-defined setting is user-defined, and in the second mode, the current value of the first user-defined setting is dynamic.
20. The method according to claim 1, wherein the one or more first criteria include criteria that are met when the first physical object is designated by the user as allowing penetration.
21. The method according to claim 1, wherein the one or more first criteria include criteria that are not met when the first physical object is designated by the user as not allowing penetration.
22. The method according to claim 1, wherein the first virtual content includes a boundary, a first part of the first virtual content, and a second part of the first virtual content that is closer to the boundary than the first part of the first virtual content, and reducing the visual prominence of the first virtual content with respect to the three-dimensional environment includes: Reducing the visual prominence of the first part of the first virtual content by a first amount with respect to the three-dimensional environment; And reducing the visual prominence of the second part of the first virtual content by a second amount greater than the first amount with respect to the three-dimensional environment.
23. The method according to claim 1, the method further comprising: In response to detecting the first physical object in the first part of the physical environment, and based on determining that the current value of the first user-defined setting is set to not allow penetration of physical objects that meet the one or more first criteria: Based on determining that the first physical object at the first location meets one or more second criteria, reducing the visual prominence of the first virtual content relative to the three-dimensional environment, where the one or more second criteria are safety-related.
24. The method according to claim 1, wherein when the first physical object is in the first part of the physical environment, the computer system is generating audio output associated with the three-dimensional environment, and the method further includes: In response to detecting the first physical object in the first part of the physical environment: Based on determining that the first physical object at the first location meets one or more first criteria and the current value of the first user-defined setting is set to allow penetration of physical objects that meet the one or more first criteria, Reducing the magnitude of the audio output associated with the three-dimensional environment.
25. The method according to claim 1, wherein reducing the visual prominence of the first virtual content relative to the three-dimensional environment includes one or more of the following: dimming, blurring, or stopping the display of one or more portions of the first virtual content.
26. The method according to claim 1, wherein the one or more first criteria include criteria that are not met when the first virtual content is system virtual content.
27. A computer system that communicates with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; And One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 1 to 26.
28. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, cause the computer system to perform the method according to any one of claims 1 to 26.