Device, method, and graphical user interface for displaying presentation environment for presentation application
By improving the computer system interface and methods, the problem of low interaction efficiency in virtual/augmented reality environments has been solved, achieving more efficient user interaction and energy saving, and extending battery life.
Patent Information
- Application Number
- CN202480037451.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-04
- Filing Date
- 2024-05-31
- Publication Date
- 2026-02-24
AI Technical Summary
Existing methods for interacting with virtual/augmented reality environments are inefficient, cumbersome, and error-prone, resulting in high energy consumption of computer systems, especially shortened battery life in battery-powered devices.
By improving computer system interfaces and methods, reducing the amount and nature of user input, providing more intuitive interaction methods, and combining visual feedback with environment-locked virtual objects, the user interface is optimized to improve efficiency and save energy.
It improves user interaction efficiency, reduces device power consumption, extends battery life, and enhances device operability and user experience.
Smart Images

Figure CN121569265A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 506,126, filed June 4, 2023, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field
[0003] This disclosure relates in its entirety to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Technology
[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices used in computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with the virtual / augmented reality environment. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention
[0005] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where manipulating virtual objects is complex, tedious, and error-prone, impose a significant cognitive burden on users and detract from the immersive experience of virtual / augmented reality environments. Furthermore, these methods take longer than necessary, thus wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0006] Therefore, computer systems with improved methods and interfaces are needed to provide users with computer-generated experiences, making user interaction with the computer system more efficient and intuitive. Such methods and interfaces optionally complement or replace conventional methods for providing users with extended reality experiences. By helping users understand the relationship between the inputs provided and the device's responses to those inputs, such methods and interfaces reduce the quantity, extent, and / or nature of user input, thus creating a more efficient human-computer interface.
[0007] The disclosed system reduces or eliminates the aforementioned defects and other problems associated with the user interface of a computer system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, tablet, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to display generation components, the computer system also has one or more output devices, including one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, and a program or instruction set stored in memory for performing multiple functions. In some implementations, the user interacts with the GUI through touch and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some implementations, functions performed through interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] Electronic devices with improved methods and interfaces are needed to interact with 3D environments. Such methods and interfaces can complement or replace conventional methods for interacting with 3D environments. They reduce the amount, extent, and / or nature of user input, resulting in more efficient human-computer interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging sessions.
[0009] In some implementations, the computer system displays one or more virtual environments associated with the demonstration application. In some implementations, the computer system initially displays the virtual demonstration in a first three-dimensional environment. In some implementations, and in response to receiving an instruction to change the displayed three-dimensional environment to a second three-dimensional environment, the computer system displays the second three-dimensional environment and also modifies the audio model applied to the audio emitted in the second three-dimensional environment according to an audio model different from the one used in the first three-dimensional environment. In some implementations, the computing system displays a three-dimensional representation of a virtual object represented in two dimensions on the virtual surface.
[0010] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive; in particular, many additional features and advantages will be apparent to those skilled in the art from the accompanying drawings, description, and claims. Furthermore, it should be pointed out that the language used in this specification has been chosen in principle for readability and instruction purposes, and such choice may not be necessary to depict or define the subject matter of the invention. Attached Figure Description
[0011] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals in all the drawings indicate corresponding parts.
[0012] Figure 1A This is a block diagram illustrating the operating environment of a computer system for providing XR experiences according to some implementation schemes.
[0013] Figures 1B to 1P It is used in Figure 1A Examples of computer systems that provide XR experiences in the operating environment.
[0014] Figure 2 This is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience, according to some implementation schemes.
[0015] Figure 3 This is a block diagram illustrating a display generation component of a computer system configured to provide an XR experience to a user, according to some implementation schemes.
[0016] Figure 4 This is a block diagram illustrating a hand tracking unit of a computer system configured to capture user gesture input according to some implementation schemes.
[0017] Figure 5 This is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input according to some implementation schemes.
[0018] Figure 6 This is a flowchart illustrating a flare-assisted gaze tracking pipeline according to some implementation schemes.
[0019] Figures 7A to 7J An exemplary user interface for displaying a virtual environment associated with a demonstration application is illustrated according to some implementation schemes.
[0020] Figure 8 This is a flowchart illustrating a process for displaying and demonstrating a virtual environment associated with an application, according to some implementation schemes.
[0021] Figures 9A to 9E An exemplary user interface for displaying a three-dimensional representation of a virtual object associated with a virtual presentation, according to some implementation schemes, is illustrated.
[0022] Figure 10 This is a flowchart illustrating a process for displaying a three-dimensional representation of a virtual object associated with a virtual presentation, according to some implementation schemes.
[0023] Figures 11A to 11D An exemplary audio model for demonstrating audio in a virtual environment associated with a demonstration application is illustrated according to some implementation schemes.
[0024] Figure 12 This is a flowchart illustrating a process for demonstrating audio in a virtual environment associated with a demonstration application, according to some implementation schemes. Detailed Implementation
[0025] According to some implementations, this disclosure relates to a user interface for providing extended reality (XR) experiences to users.
[0026] The systems, methods, and GUIs described in this paper improve user interface interactions with virtual / augmented reality environments in a variety of ways.
[0027] In some implementations, the computer system displays a virtual presentation associated with a demonstration application in one or more predetermined virtual environments. In some implementations, when the virtual presentation associated with the demonstration application is displayed in a first three-dimensional environment via a display generation component, the computer system receives, via one or more input devices, a first input corresponding to a request to display the virtual presentation in a second mode different from a first mode of the demonstration application, wherein the first three-dimensional environment includes a portion of the physical environment of the user of the computer system, and the virtual presentation is displayed in the first mode of the demonstration application.
[0028] In some implementations, the computer system displays a three-dimensional representation of a virtual object associated with the virtual presentation. In some implementations, when a virtual presentation associated with a presentation application is displayed at a first location in a three-dimensional environment via a display generation component, the computer system receives a first input via one or more input devices pointing to a two-dimensional representation of the first three-dimensional virtual object, wherein the virtual presentation includes a two-dimensional representation of the first three-dimensional virtual object. In response to receiving the first input, the computer system displays a three-dimensional representation of the three-dimensional virtual object at a second location in the three-dimensional environment, different from the first location in the three-dimensional environment, wherein the second location is outside the virtual presentation.
[0029] In some implementations, the computer system performs audio in a virtual environment associated with a demonstration application based on an audio model. In some implementations, when a virtual presentation associated with a demonstration application is displayed in a first three-dimensional environment via a display generation component, the computer system receives a first input via one or more input devices corresponding to a request to display the virtual presentation in a second three-dimensional environment, the virtual presentation including performing audio corresponding to the virtual presentation based on a first audio model associated with the first three-dimensional environment. In some implementations, in response to receiving the first input, the computer system displays the virtual presentation in a second three-dimensional environment, the virtual presentation including performing audio corresponding to the virtual presentation based on a second audio model, different from the first audio model, associated with the second three-dimensional environment.
[0030] Figures 1A to 6 Descriptions of example computer systems for providing XR experiences to users are provided (such as those described below with respect to methods 800, 1000 and / or 1200). Figures 7A to 7J Example techniques for displaying virtual environments associated with demonstration applications, according to some implementation schemes, are illustrated. Figure 8 A flowchart illustrating a process for displaying and demonstrating a virtual environment associated with an application, according to various implementation schemes, is described. Figures 7A to 7J The user interface in the example is used to demonstrate Figure 8 The process in. Figures 9A to 9E Example techniques for displaying three-dimensional representations of virtual objects associated with a virtual presentation, according to some implementation schemes, are illustrated. Figure 10 A flowchart illustrating a process for displaying a three-dimensional representation of a virtual object associated with a virtual presentation, according to various implementation schemes, is described. Figures 9A to 9E The user interface in the example is used to demonstrate Figure 10 The process in. Figures 11A to 11D An example audio model for demonstrating audio in a virtual environment associated with a demonstration application is illustrated according to some implementation schemes. Figure 12 A flowchart illustrating a process for demonstrating audio in a virtual environment associated with a demonstration application, according to various implementation schemes, is depicted. Figures 11A to 11D The user interface in the example is used to demonstrate Figure 12 The process in.
[0031] The processes described below enhance device operability and make the user-device interface more efficient through various technologies (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device). These technologies include providing users with improved visual feedback, reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional display controls, performing operations without further user input when a set of conditions are met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional technologies. These technologies also reduce power consumption and extend device battery life by enabling users to use the device faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These technologies also enable real-time communication, allow the use of fewer and / or less precise sensors, resulting in more compact, lighter, and cheaper devices, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption, thereby reducing the heat emitted by the device, which is particularly important for wearable devices, where excessive heat generated by a device within the operating parameters of its components can make wearing the device uncomfortable for the user.
[0032] Furthermore, in a method described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if the method requires performing a first step (if the conditions are satisfied) and a second step (if the conditions are not satisfied), those skilled in the art will know that the stated steps are repeated until both conditions are satisfied and conditions are not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to a method having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0033] In some implementation schemes, such as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted display (HMD), a monitor, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted or handheld device).
[0034] In describing XR experiences, various terms are used to distinguish several related but different environments that the user can sense and / or interact with (e.g., interacting with input detected by the computer system 101 that generates the XR experience, causing the computer system to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms:
[0035] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0036] Extended Reality: Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system may detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the properties of virtual objects in the XR environment may be done in response to a representation of physical motion (e.g., a voice command). People may use any of their senses—sight, hearing, touch, taste, and smell—to sense and / or interact with XR objects. For example, a person may sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. For example, audio objects can achieve audio transparency, which selectively introduces ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, people can sense and / or interact only with audio objects.
[0037] Examples of XR include virtual reality and mixed reality.
[0038] Virtual Reality: A virtual reality (VR) environment is a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and human heads are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of their presence within the computer-generated environment and / or through the simulation of a subset of their physical movements within the computer-generated environment.
[0039] Mixed Reality: Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments refer to simulated environments designed to incorporate sensory input from the physical environment, or its representations, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between, but not limited to, a purely physical environment as one end and a virtual reality environment as the other. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to demonstrate MR environments can track position and / or orientation relative to the physical environment, enabling virtual objects to interact with real objects (i.e., physical objects or their representations from the physical environment). For example, a system can cause motion so that virtual trees appear stationary relative to the physical ground.
[0040] Examples of mixed reality include augmented reality and augmented virtual reality.
[0041] Augmented Reality (AR): An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on a physical environment or a representation of the physical environment. For example, an electronic system used to demonstrate an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system may be configured to demonstrate virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and demonstrates the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceive the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when demonstrating the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto the physical environment, such as as a hologram or onto a physical surface, allowing a person to perceive the virtual objects superimposed on the physical environment. Augmented reality environments also refer to simulated environments in which the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system may transform one or more sensor images to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. As another example, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portions can be representative but not realistic versions of the original captured image. As yet another example, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.
[0042] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment combines one or more sensory inputs from a physical environment. Sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, but a person's face might be realistically reproduced from an image taken of a physical person. As another example, virtual objects might adopt the shape or color of a physical object imaged by one or more imaging sensors. As yet another example, virtual objects might adopt shadows that correspond to the sun's position within the physical environment.
[0043] In augmented reality, mixed reality, or virtual reality environments, a view of the three-dimensional environment is visible to the user. This view is typically visible to the user via a virtual viewport through one or more display generating components (e.g., a display providing stereoscopic content to different eyes of the same user), which has a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generating components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). The viewport and viewport boundary typically move with the movement of one or more display generating components (e.g., with the user's head for head-mounted devices, or with the user's hand for handheld devices such as tablets or smartphones). The user's viewpoint determines what is visible within the viewport. The viewpoint typically specifies the position and orientation relative to the 3D environment, and as the viewpoint moves, the view of the 3D environment also moves within the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptibly accurate view of the 3D environment that offers an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint shifts with the movement of the handheld or fixed device and / or with changes in the user's positioning relative to the handheld or fixed device (e.g., the user moves towards, away from, up, down, right, and / or left). For a device that includes a display generating component with virtual pass-through, portions of the physical environment visible (e.g., displayed and / or projected) via one or more display generating components are based on the field of view of one or more cameras communicating with the display generating component, which typically move with the movement of the display generating component (e.g., for a head-mounted device, it moves with the movement of the user's head, or for a handheld device such as a tablet or smartphone, it moves with the movement of the user's hand), because the user's viewpoint moves with the movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via one or more display generating components is updated based on the user's viewpoint (e.g., the displayed position and pose of the virtual objects are updated based on the movement of the user's viewpoint)).For a display generating component with optical transparency, portions of the physical environment visible through one or more display generating components (e.g., optically visible through one or more portions or fully transparent portions of the display generating component) are based on the user's field of view through the portion or fully transparent portion of the display generating component (e.g., for a head-mounted device, it moves with the movement of the user's head, or for a handheld device such as a tablet or smartphone, it moves with the movement of the user's hand), because the user's viewpoint moves with the movement of the user's field of view through the portion or fully transparent portion of the display generating component (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0044] In some embodiments, the representation of the physical environment (e.g., displayed via virtual passthrough or optical passthrough) may be partially or completely occluded by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally results in more virtual environment being displayed, replacing and / or occluding more physical environment, and decreasing the immersion level optionally results in less virtual environment being displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, displayed with increased transparency), and one or more third background objects are de-emphasized. In some embodiments, the level of immersion includes the associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by the computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment, optionally including the number of items of the displayed background content and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generating component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generating component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some implementations, background content includes user interfaces (e.g., user interfaces corresponding to applications generated by a computer system), virtual objects not associated with or included in the virtual environment and / or virtual content (e.g., files generated by the computer system or other user representations), and / or real objects (e.g., transparent objects representing real objects in the user's surrounding physical environment, visible such that they are displayed via display generation components and / or via transparent or semi-transparent components of the display generation components, because the computer system does not obscure / impede their visibility through the display generation components). In some implementations, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level is optionally displayed concurrently with background content, which is optionally displayed at full brightness, color, and / or semi-transparency.In some implementations, at higher immersion levels (e.g., a second immersion level above the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without concurrently displaying background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at a medium immersion level is displayed concurrently with dimmed, blurred, or otherwise de-emphasized background content. In some implementations, the visual characteristics of background objects differ between background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are stopped from being displayed. In some implementations, zero immersion or a zero immersion level corresponds to a virtual environment that is stopped from being displayed, and instead, a representation of the physical environment (optionally having one or more virtual objects, such as an application, window, or virtual 3D object) is displayed, and the representation of the physical environment is not occluded by the virtual environment. Using physical input elements to adjust immersion levels provides a quick and efficient way to adjust immersion, which enhances the operability of computer systems and makes user-device interfaces more efficient.
[0045] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same location and / or position within the user's viewpoint, the virtual object remains viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the direction forward of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); therefore, the user's viewpoint remains fixed even when the user's gaze shifts without moving the user's head. In embodiments where the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generating component of the computer system. For example, a viewpoint-locked virtual object displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint, even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of a viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an implementation where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so the virtual object is also referred to as a "head-locked virtual object".
[0046] Environment-locked visual objects: When a computer system displays a virtual object at a location and / or position within the user's viewpoint, the virtual object is environment-locked (or, "world-locked"), the location and / or position being based on a location and / or object within a three-dimensional environment (e.g., a physical or virtual environment) (e.g., selected and / or anchored to that location and / or object with reference to it). As the user's viewpoint moves, the location and / or object in the environment relative to the user's viewpoint changes, causing the environment-locked virtual object to appear at different locations and / or positions within the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user appears at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) so that the tree is now centered to the left in the user's viewpoint (e.g., the tree's position shifts in the user's viewpoint), the environment-locked virtual object locked to the tree appears centered to the left in the user's viewpoint. In other words, the position and / or orientation of an environment-locked virtual object displayed in the user's viewpoint depends on the position to which the virtual object is locked and / or the orientation and / or orientation of the object within the environment. In some implementations, the computer system uses a stationary frame of reference (e.g., a coordinate system anchored to a fixed position and / or object in the physical environment) to determine the position of the environment-locked virtual object displayed in the user's viewpoint. An environment-locked virtual object may be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object), or it may be locked to a movable part of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot), causing the virtual object to move with the viewpoint or that part of the environment to maintain a fixed relationship between the virtual object and that part of the environment.
[0047] In some implementations, environment-locked or viewpoint-locked virtual objects exhibit lazy following behavior, reducing or delaying their movement relative to the movement of a reference point they are following. In some implementations, when exhibiting lazy following behavior, the computer system intentionally delays the movement of the virtual object when movement of the reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) is detected. For example, when the reference point (e.g., that portion of the environment or the viewpoint) moves at a first rate, the virtual object is moved by the device to remain locked to the reference point, but moves at a second rate that is slower than the first rate (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up). In some implementations, when the virtual object exhibits lazy following behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold amount, such as 0 to 5 degrees or 0 cm to 50 cm). For example, when the reference point (e.g., a portion of the environment or viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a portion of the viewpoint or environment to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a portion of the viewpoint or environment to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases to above a threshold (e.g., a "lazy following" threshold), because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some implementations, maintaining a substantially fixed position of the virtual object relative to a reference point includes displaying the virtual object within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the reference point).
[0048] Hardware: Many different types of electronic systems enable humans to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection techniques that project graphic images onto a person's retina. Projection systems may also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface. In some embodiments, controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 2The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to scene 105 (e.g., physical environment). For example, the controller 110 is a local server located within scene 105. Alternatively, the controller 110 is a remote server (e.g., a cloud server, central server, etc.) located outside scene 105. In some embodiments, the controller 110 is communicatively coupled to display generation components 120 (e.g., HMD, monitor, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). For example, the controller 110 may be included within the housing (e.g., physical enclosure) of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more input devices in the input device 125, one or more output devices in the output device 155, one or more sensors in the sensor 190, and / or one or more peripheral devices in the peripheral device 195, or may share the same physical housing or support structure with one or more of the aforementioned devices.
[0049] In some embodiments, the display generation component 120 is configured to provide a user with an XR experience (e.g., at least the visual components of an XR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 3 The display generation component 120 is described in more detail. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation component 120.
[0050] According to some implementation schemes, when a user is virtually and / or physically present within scene 105, display generation component 120 provides the user with an XR experience.
[0051] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on his / her head, his / her hand, etc.). Thus, the display generating component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generating component 120 surrounds the user's field of view. In some embodiments, the display generating component 120 is a handheld device (such as a smartphone or tablet) configured to demonstrate XR content, and the user holds the device having a display facing the user's field of view and a camera facing scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generating component 120 is an XR chamber, housing, or room configured to demonstrate XR content, wherein the user does not wear or hold the display generating component 120. Many user interfaces described with reference to one type of hardware used for displaying XR content (e.g., a handheld device or a tripod-mounted device) can be implemented on another type of hardware used for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interaction with XR content triggered by an interaction occurring in the space in front of a handheld device or tripod-mounted device can be similarly implemented using an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface illustrating interaction with XR content triggered by movement of a handheld device or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented using an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).
[0052] Despite Figure 1A The relevant features of the operating environment 100 are illustrated herein, but those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the exemplary embodiments disclosed herein.
[0053] Figures 1A to 1PVarious examples of computer systems for performing methods and providing audio, visual, and / or haptic feedback as part of the user interface described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first display components 1-120a and second display components 1-120b and / or first optical modules 11.1.1-104a and second optical modules 11.1.1-104b) for displaying virtual elements and / or representations of the physical environment to a user of the computer system, the virtual elements and / or the representations of the physical environment optionally being generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216 to make it easier for a user who would otherwise use glasses or contact lenses to correct their vision to view the user interface, the one or more corrective lenses optionally being removably attached to one or more optical modules in the optical modules. While many user interfaces illustrated herein represent a single view of the user interface, user interfaces in HMDs optionally employ two optical modules (e.g., first display component 1-120a and second display component 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for display, one optical module for the user's right eye and a different optical module for the user's left eye, presenting slightly different images to the two different eyes to generate the illusion of stereoscopic depth. A single view of the user interface is typically a right-eye view or a left-eye view, and the depth effect is explained in text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display component 1-108) for displaying status information of the computer system to the user of the computer system (when the computer system is not worn) and / or to others near the computer system, the status information optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor components 1-356 and / or sensor components 1-356) for detecting information about the physical environment of the device. Figure 1I One or more sensors), which can be used (optionally with one or more illuminators, such as Figure 1IThe system combines the illuminators described herein to generate digital pass-through images, capture visual media corresponding to the physical environment (e.g., photographs and / or videos), or determine the pose (e.g., positioning and / or orientation) of physical objects and / or surfaces in the physical environment, enabling the placement of virtual objects based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand positioning and / or movement (e.g., sensor assemblies 1-356 and / or...). Figure 1I One or more sensors), which can be used (optionally with one or more illuminators, such as Figure 1I The illuminators 6-124 described herein (in combination) determine when one or more air gestures are performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I Eye-tracking and gaze-tracking sensors in the system), these sensors can be used (optionally combined with one or more lights, such as...) Figure 10The light (11.3.2-110) in the image determines attention or gaze localization and / or gaze movement, which may optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine a user's facial expressions and / or hand movements for generating a user avatar or representation, such as an anthropomorphic avatar or representation for real-time communication sessions, wherein the avatar has facial expressions, hand movements, and / or body movements detected by the user based on or similar to the device. Gaze and / or attention information may optionally be combined with hand tracking information to determine user interaction with one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132 and / or dials or buttons 1-328), knobs (e.g., first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), touchpads, touchscreens, keyboards, mice and / or other input devices. One or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328) are optionally used to perform system operations, such as recentering content in the user-visible 3D environment of the device, displaying the main user interface for launching an application, initiating a real-time communication session, or initiating the display of a virtual 3D background. Knobs or digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114, and / or dials or buttons 1-328) are optionally rotatable to adjust parameters of the visual content, such as the level of immersion of the virtual 3D environment (e.g., the extent to which the virtual content occupies the user's viewport in the 3D environment) or other parameters associated with the 3D environment and the virtual content displayed via optical modules (e.g., first display components 1-120a and second display components 1-120b and / or first optical modules 11.1.1-104a and second optical modules 11.1.1-104b).
[0054] Figure 1BExamples of head-mounted display (HMD) devices 1-100 configured to be worn by a user and provide virtual and altered / mixed reality (VR / AR) experiences are illustrated in front, top, and perspective views. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured at either end to the electronic strip assembly 1-104. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.
[0055] In at least one example, the band assembly 1-106 may include a first band 1-116 configured to wrap around the back of the user's head and a second band 1-117 configured to extend above the top of the user's head. As shown, the second band may extend between the first electronic band 1-105a and the second electronic band 1-105b of the electronic band assembly 1-104. The band assembly 1-104 and the band assembly 1-106 may be part of a fixing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.
[0056] In at least one example, the fixing mechanism includes a first electronic strip 1-105a, which includes a first proximal end 1-134 coupled to a display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite to the first proximal end 1-134. The fixing mechanism may also include a second electronic strip 1-105b, which includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite to the second proximal end 1-138. The fixing mechanism may also include a first strip 1-116 and a second strip 1-117, the first strip including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strip extending between the first electronic strip 1-105a and the second electronic strip 1-105b. Strips 1-105a to b and strip 1-116 may be coupled via a connecting mechanism or component 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between a first proximal end 1-134 and a first distal end 1-136, and a second end 1-148 coupled to the second electronic strip 1-105b between a second proximal end 1-138 and a second distal end 1-140.
[0057] In at least one example, the first electronic strip and the second electronic strips 1-105a to 1-105b comprise plastic, metal, or other structural materials forming the substantially rigid shape of the strips 1-105a to 1-105b. In at least one example, the first strip 1-116 and the second strip 1-117 are formed of an elastic flexible material (including woven textiles, rubber, etc.). The first strip 1-116 and the second strip 1-117 may be flexible enough to conform to the shape of the user's head when wearing the HMD 1-100.
[0058] In at least one example, one or more of the first electronic stripe and the second electronic stripe 1-105a to 1-105b may define an inner stripe volume and include one or more electronic components disposed within the inner stripe volume. In one example, such as Figure 1B As shown, the first electronic strip 1-105a may include electronic components 1-112. In one example, electronic components 1-112 may include a speaker. In another example, electronic components 1-112 may include computing components, such as a processor.
[0059] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is located in... Figure 1B The section marked 1-152 with dashed lines is because the display assembly 1-108 is configured to obscure the first opening 1-152 from a view when the HMD 1-100 is assembled. The housing 1-150 may also define a rearward second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover disposed in or across the front opening 1-152 to obscure the front opening 1-152 and a display screen (shown in other figures). In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of a user's face. The display screen of display unit 1-108 can be bent as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, wherein display unit 1-102 is pressed.
[0060] In at least one example, the housing 1-150 may define a first hole 1-126 between a first opening 1-152 and a second opening 1-154, and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 may be pressable through the corresponding holes 1-126 and 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a rotary dial and a pressable button. In at least one example, the first button 1-128 is a pressable and rotary dial button, and the second button 1-132 is a pressable button.
[0061] Figure 1C A rear perspective view of HMD 1-100 is illustrated. HMD 1-100 may include a light seal 1-110 extending rearwardly around the periphery of housing 1-150 of display assembly 1-108, as shown. The light seal 1-110 may be configured to extend from housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b disposed at or within a rearwardly facing second opening 1-154 defined by housing 1-150 and / or disposed within the internal volume of housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a to b may include a corresponding display screen 1-122a, 1-122b configured to project light toward the user's eyes in a rearward direction through the second opening 1-154.
[0062] In at least one example, reference Figure 1B and Figure 1C Both, the display assembly 1-108 can be a front-facing display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a to b can be configured to project light in a second rearward direction opposite to the first direction. As noted above, the light seal 1-110 can be configured to block light from outside the HMD 1-100 from reaching the user's eyes, including by means of... Figure 1B The front perspective view shows the light projected by the front display screen of the display assembly 1-108. In at least one example, the HMD 1-100 may also include a curtain 1-124 that blocks the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a to b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.
[0063] Figure 1B and Figure 1C Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1D to 1F In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1D to 1F Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.
[0064] Figure 1D An exploded view of an example HMD 1-200 including its various parts or components, separated according to the modularity and selective coupling of these components. For example, HMD 1-200 may include a strip 1-216 selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first strip and the second strips 1-205a to 1-205b are removably coupled to a display unit 1-202.
[0065] Furthermore, HMD 1-200 may include a light-sealing member 1-210 configured to be removably coupled to display unit 1-202. HMD 1-200 may also include a lens 1-218, which may be removably coupled to display unit 1-202, for example, on a first display assembly and a second display assembly including a display screen. Lens 1-218 may include a custom prescription lens configured for vision correction. As noted, in Figure 1D The exploded view shows that each component described above can be removably coupled, attached, reattached, and replaced to update the component, or replaced for different users. For example, belts such as belt 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a to b can be replaced according to the user, so that these components are customized to fit and correspond to a single user of HMD 1-200.
[0066] Figure 1D Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1B , Figure 1C and Figures 1E to 1FIn any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1B , Figure 1C and Figures 1E to 1F Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1D Examples of devices, features, components, and parts are shown.
[0067] Figure 1E An exploded view illustrating an example of a display unit 1-306 of an HMD is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0068] In at least one example, the display unit 1-306 may further include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a to 1-322b of the display unit 1-320 relative to the frame 1-350. In at least one example, the display unit 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a to 1-322b has at least one motor, such that the motor is capable of translating the display screens 1-322a to 1-322b to match the interpupillary distance of the user's eyes.
[0069] In at least one example, display unit 1-306 may include a dial or button 1-328 that is pressable relative to frame 1-350 and accessible to a user outside frame 1-350. Button 1-328 may be electrically connected to motor assembly 1-362 via a controller, such that button 1-328 can be operated by a user to cause the motor of motor assembly 1-362 to adjust the positioning of display screens 1-322a to 1-322b.
[0070] Figure 1E Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1D and Figure 1F In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1D and Figure 1F Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1E Examples of devices, features, components, and parts are shown.
[0071] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first and second corresponding display screen for interpupillary adjustment, as described above.
[0072] References in this article Figures 1B to 1E The following figures, which are referenced in this disclosure, will be used to describe the subject in more detail. Figure 1F The exploded view shows the various parts, systems, and components. Figure 1F The display unit 1-406 shown can be connected with Figures 1B to 1E The shown fixture assembly and integration includes electronic strips, belts, and other components (including light seals, connecting assemblies, etc.).
[0073] Figure 1F Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1E In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1E Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1F Examples of devices, features, components, and parts are shown.
[0074] Figure 1G An example is the front cover assembly 3-100 of the HMD device described herein (e.g., Figure 1G An exploded perspective view of the front cover assembly 3-1) of the HMD 3-100 shown or any other HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or “cover”), an adhesive layer 3-106, a display assembly 3-108 including a biconvex lens panel or array 3-110, and a structural decorative element 3-112. The adhesive layer 3-106 secures the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the decorative element 3-112. The decorative element 3-112 secures various components of the front cover assembly 3-100 to the frame or base of the HMD device.
[0075] In at least one example, such as Figure 1G As shown, the transparent cover 3-102, the protective cover 3-104, and the display assembly 3-108 including a biconvex lens array 3-110 can be bent to adapt to the curvature of a user's face. The transparent cover 3-102 and the protective cover 3-104 can be bent in two or three dimensions, for example, vertically in and out of the Z-plane along the Z direction, and horizontally in and out of the ZX-plane along the X direction. In at least one example, the display assembly 3-108 may include the biconvex lens array 3-110 and a display panel with pixels configured to project light through the protective cover 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., the horizontal direction) to adapt to the curvature of a user's face from one side (e.g., the left) to the other (e.g., the right). In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in the following figures, but may include the biconvex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to accommodate the curvature of the user's face.
[0076] In at least one example, the cover 3-104 may include a transparent or translucent material through which the display component 3-108 projects light. In one example, the cover 3-104 may include one or more opaque portions, such as opaque ink-printed portions or other opaque film portions on the back of the cover 3-104. When the HMD device is worn, the rear surface may be the surface of the cover 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the cover 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the cover 3-104 may include peripheral portions that visually conceal any components surrounding the outer periphery of the display screen of the display component 3-108. In this way, the opaque portions of the cover conceal any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the cover 3-104, including electronic components, structural components, etc.
[0077] In at least one example, the housing 3-104 may define one or more transparent aperture portions 3-120 through which sensors can transmit and receive signals. In one example, portion 3-120 is an aperture through which sensors can extend or transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion more transparent than the surrounding translucent or opaque portion of the housing, through which sensors can transmit and receive signals through the housing and via transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0078] Figure 1G Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, individually or in any combination. Figure 1G Examples of devices, features, components, and parts are shown.
[0079] Figure 1H An exploded view of an example HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102, which includes one or more sensors, cameras, projectors, etc., mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 on which one or more sensors of the sensor system 6-102 may be fixed / secured.
[0080] Figure 1I A portion of an HMD device 6-100, including a front transparent cover 6-104 and a sensor system 6-102, is illustrated. The sensor system 6-102 may include multiple different sensors, transmitters, and receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated on the front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter in system 6-102. As referenced herein, "side," "side," "lateral," "horizontal," and other similar terms refer to... Figure 1J The orientation or direction indicated by the X-axis. Terms such as "vertical," "upward," "downward," and similar terms refer to the orientation or direction indicated by... Figure 1JThe orientation or direction indicated by the Z-axis. Terms such as "frontward," "rearward," "forward," "backward," and similar terms refer to the orientation or direction indicated by the Z-axis. Figure 1J The orientation or direction indicated by the Y-axis shown.
[0081] In at least one example, a transparent cover 6-104 may define the front outer surface of an HMD device 6-100, and a sensor system 6-102, including various sensors and their components, may be positioned behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through it, including both light detected by the sensor system 6-102 and light emitted therefrom.
[0082] As noted elsewhere herein, the HMD device 6-100 may include one or more controllers, which include a processor for electrically coupling various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as displays. Furthermore, as will be shown in more detail below with reference to other accompanying drawings, various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to the HMD device 6-100. Figure 1I Various structural frame components, brackets, etc., not shown. For clarity, Figure 1I The components of the sensor system 6-102 are shown, which are not attached to or electrically coupled to other components.
[0083] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. These instructions may include, or cause the processor to execute one or more algorithms for self-correcting the angle and position of the various cameras described herein as the camera's initial position, angle, or orientation is affected by collisions or deformations due to accidental drop events or other events over time.
[0084] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. System 6-102 may include two scene cameras 6-102, respectively positioned on either side of the nose bridge or arched structure of the HMD device 6-100, such that each of the two cameras 6-106 approximately corresponds to the positioning of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and, when the HMD device 6-100 is used, provide images and content for MR video pass-through to a display screen facing the user's eyes. The scene cameras 6-106 may also be used for environment and object reconstruction.
[0085] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 that is generally pointing forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction as well as user hand and body tracking. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centrally located along the width of the HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be located above the central bridge of the nose or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction as well as hand and body tracking. In at least one example, the second depth sensor may include a LiDAR sensor.
[0086] In at least one example, the sensor system 6-102 may include a depth projector 6-112, which is typically forward-facing to project electromagnetic waves (e.g., in the form of a predetermined spot pattern) into or within the field of view of the user and / or scene camera 6-106, or into or beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a spot pattern, which are reflected from an object and back into the depth sensors (including depth sensors 6-108, 6-110) indicated above. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction, as well as hand and body tracking.
[0087] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114, whose field of view is generally directed downwards relative to the HMD device 6-100 on the Z-axis. In at least one example, the downward-facing camera 6-114 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, headset tracking, and facial image detection and creation for displaying a user's image on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward-facing camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.
[0088] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, headset tracking, and facial image detection and creation for displaying a user's image on the front display screen of the HMD device 6-100, as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. This is used for hand and body tracking, headset tracking, and facial image creation.
[0089] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views along the X-axis or in a direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, headphone tracking, and facial detection and reconstruction.
[0090] In at least one example, the sensor system 6-102 may include multiple eye-tracking and gaze-tracking sensors for determining identity, status, and the user's gaze direction during and / or prior to use. In at least one example, the eye / gaze-tracking sensor may include a nose-eye camera 6-120 positioned on either side of the user's nose and adjacent to the user's nose when wearing the HMD device 6-100. The eye / gaze sensor may also include a bottom eye camera 6-122 positioned below the respective user's eyes for capturing images of the eyes for use in facial avatar detection and creation, gaze tracking, and iris identification functions.
[0091] In at least one example, sensor system 6-102 may include an infrared illuminator 6-124 that is pointed outward from HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of sensor system 6-102. In at least one example, sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, flicker sensor 6-126 may detect the refresh rate of the overhead light to avoid display flicker. In one example, infrared illuminator 6-124 may include a light-emitting diode and may be specifically designed for low-light environments to illuminate a user's hands and other objects in low light for detection by the infrared sensors of sensor system 6-102.
[0092] In at least one example, multiple sensors (including scene camera 6-106, downward camera 6-114, chin camera 6-116, side camera 6-118, depth projector 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination, thereby improving the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the above-described and Figure 1I The downward-facing camera 6-114, the jaw camera 6-116, and the side camera 6-118 shown can be wide-angle cameras capable of operating in both the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, and 6-118 can operate solely in black-and-white light detection to simplify image processing and achieve sensitivity.
[0093] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1J to 1L In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1J to 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1I Examples of devices, features, components, and parts are shown.
[0094] Figure 1JA lower perspective view of an example HMD 6-200 including a cover or shield 6-204 fixed to a frame 6-230 is shown. In at least one example, a sensor 6-203 of a sensor system 6-202 may be disposed around the periphery of the HMD 6-200 such that the sensor 6-203 is disposed outwardly around the periphery of the display area or region 6-232 so as not to obstruct the view of the displayed light. In at least one example, the sensor may be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing light to pass back and forth through the shield 6-204 by the sensor and the projector. In at least one example, an opaque ink or other opaque material or film / layer may be disposed on the shield 6-204 around the display area 6-232 to conceal components of the HMD 6-200 outside the display area 6-232 rather than through a transparent portion defined by the opaque portion through which the sensor and the projector transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area surrounding the periphery of the display and the shield 6-204.
[0095] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 transmits and receives signals. In the illustrated examples, the sensor 6-203 of the sensor system 6-202, which transmits and receives signals through the shield 6-204, or more specifically through the transparent area 6-209 defined by the opaque portion 6-207 of the shield 6-204, may include... Figure 1I The examples illustrate those same or similar sensors, such as depth sensors 6-108 and 6-110, depth projector 6-112, first scene camera and second scene camera 6-106, first downward camera and second downward camera 6-114, first side camera and second side camera 6-118, and first infrared illuminator and second infrared illuminator 6-124. These sensors also... Figure 1K and Figure 1L The example is shown. Other sensors, sensor types, number of sensors, and their relative positioning can be included in one or more other examples of the HMD.
[0096] Figure 1J Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1I and Figures 1K to 1LIn any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1I and Figures 1K to 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1J Examples of devices, features, components, and parts are shown.
[0097] Figure 1K A front view of a portion of an example of an HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330, is shown. Figure 1K The examples shown do not include a front cover or shield to illustrate brackets 6-336 and 6-338. For example, Figure 1J The shield 6-204 shown includes an opaque portion 6-207 that visually covers / blocks the view of anything outside the display / display area 6-334 (e.g., radially / peripherally outside the display / display area), including the sensor 6-303 and the bracket 6-338.
[0098] In at least one example, various sensors of sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, scene camera 6-306 includes strict tolerances for its angle relative to each other. For example, the tolerance for the mounting angle between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such strict tolerances, in one example, scene camera 6-306 may be mounted to bracket 6-338 instead of a housing. The bracket may include a cantilever on which scene camera 6-306 and other sensors of sensor system 6-302 may be mounted to maintain their positioning and orientation in the event of a drop event caused by a user that results in any deformation of other brackets 6-226, housing 6-330, and / or housing.
[0099] Figure 1K Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1J and Figure 1L In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1I to 1J and Figure 1L Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1K Examples of devices, features, components, and parts are shown.
[0100] Figure 1LA bottom view illustrating an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is shown. The sensor system 6-402 is compatible with the above and other parts of this document (including references). Figures 1I to 1K Other sensor systems described are similar. In at least one example, the jaw camera 6-416 may be oriented downwards to capture images of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to a frame or housing 6-430 or one or more internal brackets directly coupled to the frame or housing 6-430 shown. The frame or housing 6-430 may include one or more holes / openings 6-415 through which the jaw camera 6-416 transmits and receives signals.
[0101] Figure 1L Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1K In any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1I to 1K Any of the features, components and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1L Examples of devices, features, components, and parts are shown.
[0102] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated. This IPD adjustment system includes a first optical module and a second optical module 11.1.1-104a-11.1.1-104b that are slidably engaged / coupled to corresponding guide rods 11.1.1-108a-11.1.1-108b and motors 11.1.1-110a-11.1-110b of the left and right adjustment subsystems 11.1.1-106a-11.1-106b. The IPD adjustment system 11.1.1-102 is coupled to a bracket 11.1.1-112 and includes buttons 11.1.1-114 that are electrically in communication with the motors 11.1.1-110a-11.1.1-110b. In at least one example, buttons 11.1.1-114 can be electrically communicated with the first motor and the second motors 11.1.1-110a to 11.1.1-110b via a processor or other circuit components to activate the first motor and the second motors 11.1.1-110a to 11.1.1-110b and respectively cause the first optical module and the second optical modules 11.1.1-104a to 11.1.1-104b to change their positions relative to each other.
[0103] In at least one example, the first and second optical modules 11.1.1-104a to 11.1.1-104b may include corresponding display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, a user-operable (e.g., pressing and / or rotating) button 11.1.1-114 activates a positioning adjustment of the optical modules 11.1.1-104a to 11.1.1-104b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a to 11.1.1-104b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD, such that the optical modules 11.1.1-104a to 11.1.1-104b can be adjusted to match the IPD.
[0104] In one example, a user can manipulate button 11.1.1-114 to induce automatic positioning adjustment of the first and second optical modules 11.1.1-104a to 11.1.1-104b. In another example, a user can manipulate button 11.1.1-114 to induce manual adjustment, causing the optical modules 11.1.1-104a to 11.1.1-104b to move further or closer (e.g., when the user rotates button 11.1.1-114 in one way or another) until the user visually aligns it with their own IPD. In one example, manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a to 11.1.1-104b via motors 11.1.1-110a to 11.1.1-110b is supplied by a power source. In one example, the adjustment and movement of optical modules 11.1.1-104a to 11.1.1-104b via manipulation buttons 11.1.1-114 are mechanically actuated via movement buttons 11.1.1-114.
[0105] Figure 1M Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and parts shown and described herein in any other illustrated figures. Similarly, any of the features, components, and / or parts shown and described herein with reference to any other illustrated figures (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and / or parts shown and / or described herein. Figure 1M Examples of devices, features, components, and parts are shown.
[0106] Figure 1NA front perspective view of a portion of HMD 11.1.2-100 is shown, including an outer structural frame 11.1.2-102 defining a first hole 11.1.2-106a and a second hole 11.1.2-106b, and an inner or intermediate structural frame 11.1.2-104. Holes 11.1.2-106a to 11.1.2-106b are located in... Figure 1N The holes 11.1.2-106a to 11.1.2-106b are shown in dashed lines because viewing the HMD 11.1.2-100 may be obstructed by one or more other components coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second holes 11.1.2-106a to 11.1.2-106b.
[0107] Mounting brackets 11.1.2-108 may include intermediate or central portions 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portions 11.1.2-109 may not be the geometric center or middle of the brackets 11.1.2-108. Instead, the intermediate / central portions 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portions 11.1.2-109. In at least one example, mounting bracket 108 includes first cantilever arms 11.1.2-112 and second cantilever arms 11.1.2-114 extending away from the intermediate portions 11.1.2-109 of the mounting brackets 11.1.2-108 coupled to the inner frame 11.1.2-104.
[0108] like Figure 1NAs shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to adapt to the user's nose when the user wears the HMD 11.1.2-100. This curved geometry may be referred to as the bridge of the nose 11.1.2-111 and is centrally located on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between holes 11.1.2-106a to 11.1.2-106b, such that the cantilever 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the central portion 11.1.2-109 to complement the bridge of the nose 11.1.2-111 geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to adapt to the user's nose, as noted above. The geometry of the bridge 11.1.2-111 adapts to the nose because the bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, above, and around.
[0109] The first cantilever 11.1.2-112 may extend in a first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108, and the second cantilever 11.1.2-114 may extend in a second direction opposite to the first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as “cantilever” or “cantilever” arms because each arm 11.1.2-112, 11.1.2-114 includes free distal ends 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, arms 11.1.2-112 and 11.1.2-114 extend from the middle section 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102 and 11.1.2-104 are not attached.
[0110] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a to 11.1.2-110f. Each of the plurality of sensors 11.1.2-110a to 11.1.2-110f may include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a to 11.1.2-110f may be used for object recognition in three-dimensional space, making it important to maintain the precise relative positioning of two or more of the plurality of sensors 11.1.2-110a to 11.1.2-110f. The cantilever nature of the mounting bracket 11.1.2-108 protects the sensors 11.1.2-110a to 11.1.2-110f from damage and displacement in the event of accidental drops by the user. Because sensors 11.1.2-110a to 11.1.2-110f cantilevered on arms 11.1.2-112 and 11.1.2-114 of mounting bracket 11.1.2-108, stress and deformation of the internal frame and / or external frame 11.1.2-104 and 11.1.2-102 are not transmitted to the cantilever 11.1.2-112 and 11.1.2-114, and therefore do not affect the relative position of sensors 11.1.2-110a to 11.1.2-110f coupled to / mounted to mounting bracket 11.1.2-108.
[0111] Figure 1N Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in any other example of the devices, features, components, and other examples described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 1N Examples of devices, features, components, and parts are shown.
[0112] Figure 10 An example of optical modules 11.3.2-100 for use in electronic devices, such as HMDs, including the HDM devices described herein, is illustrated. As shown in one or more other examples described herein, optical module 11.3.2-100 may be one of two optical modules within an HMD, wherein each optical module is aligned to project light toward a user's eye. In this way, a first optical module may project light toward a user's first eye via a display screen, and a second optical module of the same device may project light toward a user's second eye via another display screen.
[0113] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a tube or optical module tube. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the user's eyes when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0114] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eye during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the camera 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward the user's eyes when the HMD is worn. The individual lights 11.3.2-110 in the light strips 11.3.2-108 may be spaced apart around the light strips 11.3.2-108, and are therefore uniformly or non-uniformly spaced around the display 11.3.2-104 at various locations on the light strips 11.3.2-108 and around the display 11.3.2-104.
[0115] In at least one example, housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light onto the user's eyes through the viewing opening 11.3.2-101. In one example, camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.
[0116] As pointed out above, Figure 10 Each of the components and features of the optical modules 11.3.2-100 shown can be replicated in another (e.g., a second) optical module set up with the HMD to interact with the user’s other eye (e.g., project light and capture images).
[0117] Figure 10 Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1P Any other example of the device, feature, component, and part shown or otherwise described herein. Similarly, refer to... Figure 1P Any of the features, components, and / or parts shown, described, or otherwise described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 10 Examples of devices, features, components, and parts are shown.
[0118] Figure 1P A cross-sectional view of an example optical module 11.3.2-200 is shown, which includes a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. Channels 11.3.2-212 and 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of an HMD device to allow the optical module 11.3.2-200 to be adjusted and positioned relative to the user's eye to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guide rod to secure the optical module 11.3.2-200 in the appropriate position within the HMD.
[0119] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and positioned between the display assembly 11.3.2-204 and the user's eye when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eye. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, lenses 11.3.2-216 are positioned above light strips 11.3.2-208 and one or more eye-tracking cameras 11.3.2-206, such that cameras 11.3.2-206 are configured to capture an image of a user's eye through lenses 11.3.2-216, and light strips 11.3.2-208 include lamps configured to project light onto the user's eye through lenses 11.3.2-216 during use.
[0120] Figure 1P Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in any other example of the devices, features, components, and parts described herein and in any other example. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 1P Examples of devices, features, components, and parts are shown.
[0121] Figure 2This is a block diagram of an example controller 110 according to some implementation schemes. Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the implementation schemes disclosed herein. Therefore, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs) and / or processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0122] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, and / or one or more displays.
[0123] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures, or subsets thereof, including optional operating system 230 and XR experience module 240.
[0124] Operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., single XR experiences for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0125] In some implementations, the data acquisition unit 241 is configured to acquire data from... Figure 1A The data acquisition unit 241 includes at least the display generation component 120, and optionally acquires data (e.g., presentation data, interactive data, sensor data, location data, etc.) from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0126] In some implementations, the tracking unit 242 is configured to map scene 105, and the tracking at least shows the generated component 120 relative to... Figure 1A The scenario 105 and optionally the location / position relative to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the location / position of one or more portions of the user's hand, and / or the location / position of one or more portions of the user's hand relative to... Figure 1A Scenario 105, movement relative to the display generating component 120 and / or relative to a coordinate system (defined relative to the user's hand). The following refers to movement relative to... Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the user's gaze (or more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the positioning or movement of XR content displayed via display generation component 120. The following description is relative to... Figure 5 The eye-tracking unit 243 is described in more detail.
[0127] In some implementations, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripheral device 195. To this end, in various implementations, coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0128] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0129] Although the data acquisition unit 241, the tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 may reside in a separate computing device.
[0130] also, Figure 2 This is used more as a functional description of various features that can exist in a particular specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0131] Figure 3This is a block diagram illustrating an example of generating component 120 according to some embodiments. Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the display generating component 120 (e.g., HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, and / or processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.
[0132] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, and / or one or more depth sensors (e.g., structured light and / or time-of-flight, etc.).
[0133] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, display generation component 120 (e.g., HMD) includes a single XR display. Alternatively, display generation component 120 may include XR displays for each of the user's eyes. In some embodiments, the one or more XR displays 312 are capable of displaying MR and VR content. In some embodiments, the one or more XR displays 312 are capable of displaying MR or VR content.
[0134] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to the scene the user would see in the absence of a display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0135] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and XR demonstration module 340.
[0136] Operating system 330 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR demonstration module 340 is configured to demonstrate XR content to a user via one or more XR displays 312. Therefore, in various embodiments, XR demonstration module 340 includes a data acquisition unit 342, an XR demonstration unit 344, an XR image generation unit 346, and a data transmission unit 348.
[0137] In some implementations, the data acquisition unit 342 is configured to acquire data from at least... Figure 1A The controller 110 acquires data (e.g., demonstration data, interaction data, sensor data, location data, etc.). To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0138] In some implementations, the XR presentation module 344 is configured to present XR content via one or more XR displays 312. To this end, in various implementations, the XR presentation unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0139] In some implementations, the XR graph generation unit 346 is configured to generate XR graphs based on media content data (e.g., 3D graphs of mixed reality scenes or graphs in which computer-generated objects can be placed to generate extended reality physical environments). To this end, in various implementations, the XR graph generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0140] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., demonstration data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 348 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0141] Although the data acquisition unit 342, XR demonstration unit 344, XR image generation unit 346, and data transmission unit 348 are shown residing in a single device (e.g., Figure 1A The display generation unit 120 is located on the display generation unit, but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR demonstration unit 344, the XR image generation unit 346, and the data transmission unit 348 may be located in a separate computing device.
[0142] also, Figure 3 This is used more as a functional description of various features that can exist in a particular specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 3 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0143] Figure 4 This is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) by hand tracking unit 244 ( Figure 2 To control and track the location / position of one or more parts of a user's hand, and / or the location of one or more parts of the user's hand relative to the user's hand. Figure 1A The scenario 105 refers to movement relative to a portion of the physical environment surrounding the user, relative to display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined relative to the user's hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0144] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures at least three-dimensional scene information including the human user's hand 406. The image sensor 404 captures images of the hand at sufficient resolution to allow differentiation of the fingers and their corresponding positions. The image sensor 404 typically captures images of other parts of the user's body, or possibly all parts of the body, and may have scaling capabilities or a dedicated sensor with increased magnification to capture images of the hand at the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of scene 105, or serves as the image sensor for capturing the physical environment of scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a manner that uses the field of view of the image sensor 404 or a portion thereof to define an interaction space in which hand movements captured by the image sensor are considered input to the controller 110.
[0145] In some implementations, image sensor 404 outputs a sequence of frames containing 3D image data (and, in addition, possibly color image data) to controller 110, which extracts high-level information from the image data. This high-level information is typically provided via an application programming interface (API) to an application running on the controller, which in turn drives display generation component 120. For example, a user can interact with software running on controller 110 by moving his hand 406 and changing his hand gestures.
[0146] In some embodiments, image sensor 404 projects a speckle pattern onto a scene including hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) via triangulation based on the lateral offset of the specks in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a certain distance from image sensor 404. In this disclosure, it is assumed that image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand-tracking device) may use other 3D mapping methods, such as stereo imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.
[0147] In some implementations, hand tracking device 140 captures and processes time-series depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D map data to extract image block descriptors of the hand from these depth maps. The software may match these descriptors with image block descriptors stored in database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D position of the user's hand joints and fingertips.
[0148] The software can also analyze the trajectories of the hand and / or fingers across multiple frames in the sequence to identify gestures. The pose estimation function described herein can be alternated with the motion tracking function, such that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find pose changes occurring in the remaining frames. Pose, motion, and gesture information is provided to an application running on controller 110 via the aforementioned API. The program can, for example, move and modify the image presented on display generation component 120 in response to the pose and / or gesture information, or perform other functions.
[0149] In some implementations, gestures include air gestures. An air gesture is a gesture detected without the user touching an input element that is part of the device (e.g., computer system 101, one or more input devices 125 and / or hand tracking device 140) (or independent of an input element that is part of the device) and based on the detected movement of a part of the user's body (e.g., head, one or two arms, one or two hands, one or more fingers and / or one or two legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to one of the user's fingers or a part of the user's other hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0150] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures performed by the movement of a user's fingers relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, air gestures are detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and are based on the detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's other hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0151] In some implementations where the input gesture is an air gesture (e.g., where the input device provides information to the computer system about which user interface element is the target of the user input in the absence of physical contact, such as contact with a user interface element displayed on a touchscreen, or contact with a mouse or touchpad to move the cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Therefore, in implementations involving air gestures, for example, the input gesture is combined with (e.g., simultaneously) movement of the user's fingers and / or hand to detect attention (e.g., gaze) toward a user interface element to perform pinch and / or tap input, as described in more detail below.
[0152] In some implementations, input gestures directed to a user interface object are performed, either directly or indirectly, by referencing the user interface object. For example, user input is performed directly on the user interface object based on the location of the user's hand corresponding to the positioning of the user interface object in a three-dimensional environment (e.g., determined based on the user's current viewpoint). In some implementations, when user attention to the user interface object (e.g., gazing) is detected, input gestures are performed indirectly on the user interface object based on the user's hand not being positioned at a location corresponding to the positioning of the user interface object in the three-dimensional environment at the time the user performs the input gesture. For example, for direct input gestures, the user can direct their input to the user interface object by initiating a gesture at or near the displayed positioning of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 cm and 5 cm, such as measured from the outer edge or center of an option). For indirect input gestures, the user can direct their input toward the user interface object by focusing on it (e.g., by looking at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location that the computer system can detect) (e.g., at a location that does not correspond to the displayed location of the user interface object).
[0153] In some implementations, the input gestures (e.g., air gestures) used in the various examples and implementations described herein include pinch input and tap input for interacting with a virtual or mixed reality environment. For example, the pinch input and tap input described below are performed as air gestures.
[0154] In some implementations, pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 to 1 second) interruption of contact. A long pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact is detected. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some implementations, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively with each other immediately (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts the contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).
[0155] In some embodiments, pinch and drag gestures as air gestures include pinch gestures (e.g., pinching gestures or long pinch gestures) performed in conjunction with (e.g., following) drag inputs that change the user's hand position from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opening two or more of their fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together to touch each other and uses the drag gesture to move the same hand to the second position in the air). In some implementations, pinch input is performed by the user's first hand, and drag input is performed by the user's second hand (e.g., while the user continues pinch input using their first hand, their second hand moves in the air from a first position to a second position). In some implementations, input gestures as air gestures include inputs performed using both of the user's hands (e.g., pinch and / or tap inputs). For example, input gestures include two (e.g., more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., pinch input, long pinch input, or pinch and drag input) is performed using the user's first hand, and a second pinch input is performed using the other hand (e.g., the second hand among the user's two hands).
[0156] In some embodiments, a tap input performed as an air gesture (e.g., pointing at a user interface element) includes movement of a user's finger toward the user interface element, movement of a user's hand toward the user interface element (optionally, where the user's finger extends toward the user interface element), downward movement of a user's finger (e.g., mimicking a mouse click or a tap on a touchscreen), or other predefined movements of the user's hand. In some embodiments, the tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the finger or hand moving away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on changes in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of finger or hand movement, and / or a reversal of the acceleration direction of finger or hand movement).
[0157] In some implementations, the user's attention is determined to be directed to a portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment (optionally, no other conditions are required). In some implementations, the user's attention is determined to be directed to that portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment using one or more additional conditions, such as requiring the gaze to be directed to that portion of the 3D environment for at least a threshold duration (e.g., dwell time) and / or requiring the gaze to be directed to that portion of the 3D environment when the user's viewpoint is within a distance threshold from that portion of the 3D environment, so that the device determines that the user's attention is directed to that portion of the 3D environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the 3D environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0158] In some implementations, the detection of the readiness configuration of a user or a portion of a user is performed by a computer system. The detection of the hand's readiness configuration is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the readiness of the hand is determined based on whether it has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grasping gesture, or a pre-tap with one or more fingers extended and the back of the hand facing the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15cm, 20cm, 25cm, 30cm, or 50cm from the body), and / or whether the hand has moved in a particular manner (e.g., moving towards the area in front of the user above the user's waist and below the user's head, or moving away from the user's body or legs). In some implementations, the readiness state is used to determine whether an interactive element of the user interface responds to attentional (e.g., gaze) input.
[0159] In scenarios where input is described using reference to air gestures, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such gestures. Optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the spatial positioning of the hardware input device, and the positioning and / or movement of the hardware input device can be used in place of the positioning and / or movement of the one or two hands corresponding to the air gesture. In scenarios where input is described using reference to air gestures, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such gestures. User input can be detected using controls contained in hardware input devices, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, covers of one or two hands or one or two fingers that can detect the positioning or positional change of parts of the hand and / or fingers relative to each other, relative to the user's body and / or relative to the user's physical environment, and / or other hardware input device controls, wherein user input using controls contained in the hardware input device replaces hand and / or finger gestures such as air taps or air pinches in corresponding air gestures. For example, a selection input described as being performed using air tap or air pinch input can alternatively be detected using button presses, taps on touch-sensitive surfaces, presses on pressure-sensitive surfaces, or other hardware inputs. As another example, motion input described as being performed using air pinch and drag can be optionally detected based on interaction with hardware input controls, such as pressing and holding a button, touching a touch on a touch-sensitive surface, pressing a pressure-sensitive surface, or other hardware input following movement of a hardware input device (e.g., a hand associated with the hardware input device) through space. Similarly, two-handed input involving movement of hands relative to each other can be performed using an air gesture and a hardware input device in which the hand is not performing the air gesture, two hardware input devices held in different hands, or two air gestures performed by different hands using air gestures and / or inputs detected by one or more hardware input devices described above.
[0160] In some embodiments, the software may be downloaded to controller 110 electronically, for example, via a network, or alternatively, may be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, database 408 is also stored in memory associated with controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware, such as custom or semi-custom integrated circuits or programmable digital signal processors (DSPs). Although in Figure 4The controller 110 is shown, but for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing function of the image sensor 404 may also be integrated into a computer or other computerized device controlled by the sensor output.
[0161] Figure 4 Also included is a schematic diagram of a depth map 410 captured by image sensor 404 according to some embodiments. As explained above, the depth map comprises a matrix of pixels with corresponding depth values. Pixel 412 corresponding to the hand 406 has been segmented from the background and wrist in the map. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where gray shadows become darker as depth increases. Controller 110 processes these depth values to identify and segment components of the image that have human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and frame-to-frame motion from the depth map sequence.
[0162] Figure 4 The controller 110 also schematically illustrates, according to some embodiments, the hand skeleton 414 ultimately extracted from the depth map 410 of the hand 406. Figure 4 In this configuration, the hand skeleton 414 is superimposed on the hand background 416, which has already been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally those on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, fingertips, the center of the palm, the end of the hand that connects to the wrist, etc.), are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.
[0163] Figure 5 An eye-tracking device 130 is illustrated. Figure 1A Example implementation of ). In some implementations, the eye-tracking device 130 consists of an eye-tracking unit 243 ( Figure 2The eye-tracking device 130 controls the positioning and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as a headset, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye-tracking device 130 is optionally a separate device from the handheld device or XR room. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or not head-mounted. In some embodiments, the eye-tracking device 130 is not a head-mounted device and is optionally used in conjunction with head-mounted display generation components. In some embodiments, the eye-tracking device 130 is not a head-mounted device and is optionally part of non-head-mounted display generation components.
[0164] In some embodiments, the display generation unit 120 uses display mechanisms (e.g., a left near-eye display panel and a right near-eye display panel) to display frames including left and right images in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation unit may include left and right optical lenses (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation unit may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation unit may have a transparent or semi-transparent display on which virtual objects are displayed, allowing the user to view the physical environment directly through the transparent or semi-transparent display. In some embodiments, the display generation unit projects virtual objects onto the physical environment. The virtual objects may, for example, be projected onto a physical surface or as holograms, allowing an individual to observe virtual objects superimposed on the physical environment using the system. In this case, separate display panels and image frames for the left and right eyes may not be necessary.
[0165] like Figure 5As shown, in some embodiments, eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an array or ring of IR or NIR light sources, such as LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye-tracking camera may be pointed at the user's eyes to receive IR or NIR light reflected directly from the eyes, or alternatively, it may be pointed at "hot" mirrors located between the user's eyes and the display panel, which reflect the IR or NIR light from the eyes back to the eye-tracking camera while allowing visible light to pass through. Eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps), analyzes these images to generate gaze tracking information, and transmits the gaze tracking information to controller 110. In some embodiments, the user's two eyes are tracked separately using corresponding eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked using corresponding eye-tracking cameras and illumination sources.
[0166] In some embodiments, a device-specific calibration procedure is used to calibrate the eye-tracking device 130 to determine parameters of the eye-tracking device for a specific operating environment 100, such as the 3D geometry and parameters of the LEDs, camera, thermal mirror (if present), eye lenses, and display screen. The device-specific calibration procedure can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration procedure can be an automated or manual calibration procedure. According to some embodiments, a user-specific calibration procedure may include estimation of eye parameters for a specific user, such as pupil position, foveal position, optical axis, visual axis, interocular distance, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, a flash-assisted method can be used to process the images captured by the eye-tracking camera to determine the current visual axis and the user's gaze point relative to the display.
[0167] like Figure 5As shown, the eye-tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system. The gaze tracking system includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye-tracking camera 540 may be pointed toward a mirror 550 located between the user's eye 592 and a display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, projector, etc.). These mirrors reflect the IR or NIR light from the eye 592 while allowing visible light to pass through. Figure 5 (as shown in the top portion), or alternatively, it can be pointed towards the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion), Figure 5 (As shown in the bottom part).
[0168] In some implementations, controller 110 renders AR or VR frames 562 (e.g., left and right frames for the left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye-tracking camera 540 for various purposes, such as processing frame 562 for display. Controller 110 optionally estimates the user's gaze point on display 510 based on the gaze tracking input 542 obtained from eye-tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated based on gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0169] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, controller 110 may generate virtual content at a higher resolution in the concave area determined according to the user's current gaze direction than in the peripheral area. As another example, the controller may position or move virtual content in the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content in the view based at least partially on the user's current gaze direction. As another example use case in AR applications, controller 110 may guide an external camera used to capture the physical environment of an XR experience to focus in the determined direction. The external camera's autofocus mechanism may then focus on an object or surface in the environment that the user is currently looking at on display 510. As another example use case, eye lens 520 may be a focusable lens, and the controller uses gaze tracking information to adjust the focus of eye lens 520 so that the virtual object the user is currently looking at has appropriate convergence / divergence to match the convergence of the user's eyes 592. The controller 110 can use gaze tracking information to guide the eye lens 520 to adjust its focus so that the nearby object that the user is looking at appears at the correct distance.
[0170] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lens 520), an eye-tracking camera (e.g., eye-tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)) mounted in a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each lens in the head-mounted device, such as... Figure 5 As shown. In some embodiments, as an example, eight lighting sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer lighting sources 530 may be used, and other arrangements and positions of the lighting sources 530 may be used.
[0171] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the positions and angles of the eye-tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at different wavelengths (e.g., 940 nm) may be used on each side of the user's face.
[0172] like Figure 5 The illustrated gaze tracking system implementation can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0173] Figure 6 Examples of flash-assisted gaze tracking pipelines according to some embodiments are illustrated. In some embodiments, the gaze tracking pipeline uses a flash-assisted gaze tracking system (e.g., such as...) Figure 1A and Figure 5 The illustrated eye-tracking device 130 is used to implement this. The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and flash in the current frame. When not in tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in tracking state.
[0174] like Figure 6 As shown, the gaze-tracking camera captures left and right images of the user's left and right eyes. The captured images are then fed into a gaze-tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze-tracking system can continue capturing images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments or under certain conditions, not all captured frames are processed by the pipeline.
[0175] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, the image is analyzed to detect the user's pupils and flashes of light, as indicated at 620. At 630, if the pupils and flashes of light are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eyes.
[0176] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flashes in part based on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the detected pupils and flashes in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result is credible. For example, the result may be checked to determine whether a sufficient number of pupils and flashes used for gaze estimation were successfully tracked or detected in the current frame. At 650, if the result is not credible, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is credible, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.
[0177] Figure 6 This is intended as an example of an eye-tracking technology that can be used in a particular specific implementation. As will be recognized by those skilled in the art, in a computer system 101 for providing an XR experience to a user, other eye-tracking technologies that are currently available or will be developed in the future may be used to replace or in combination with the flash-assisted eye-tracking technology described herein, depending on the various implementations.
[0178] In some implementations, a portion of the captured real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0179] Therefore, this description describes some implementations of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and virtual objects. For example, a three-dimensional environment optionally includes a representation of a table existing in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of a computer system or passively displayed via a transparent or semi-transparent display of a computer system). As previously described, the three-dimensional environment is optionally a mixed reality system in which the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in a three-dimensional environment to appear as if the virtual objects exist in the real world (e.g., the physical environment) by placing virtual objects in the three-dimensional environment at corresponding locations in the real world that have corresponding positions in the three-dimensional environment. For example, the computer system optionally displays a vase such that the vase appears as if a real vase were placed on top of a table in the physical environment. In some implementations, a corresponding location in the three-dimensional environment has a corresponding location in the physical environment. Therefore, when a computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., such as at or near a user's hand or at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object were at or near a physical object in the physical environment (e.g., the virtual object is displayed in the three-dimensional environment at a location that would be displayed in the physical environment if the virtual object were a real object at that specific location).
[0180] In some implementations, real-world objects that exist in a physical environment and are displayed in a 3D environment (e.g., and / or visible via display-generated components) can interact with virtual objects that exist only in the 3D environment. For example, the 3D environment may include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.
[0181] In a three-dimensional environment (e.g., a real environment, a virtual environment, or a hybrid environment including both real and virtual objects), an object is sometimes referred to as having depth or simulated depth, or as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension other than height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to a user's position or viewpoint, in which case the depth dimension varies based on the user's position and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the user's location relative to a surface of the environment (e.g., the surface of the environment's floor or ground), objects further away from the user along lines extending parallel to the surface are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder extending from the user's head toward the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., a direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), objects further away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from a line extending from and parallel to the user's viewpoint (e.g., defining depth in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere extending outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application displaying application and / or system content), where the user interface container has a height and / or width, and depth is a dimension orthogonal to the height and / or width of the user interface container. In some implementations, when the depth is defined relative to the user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container are typically orthogonal or substantially orthogonal to a straight line extending from the user's location (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some implementations, when the depth is defined relative to the user interface container, the depth of an object relative to the user interface container refers to the object's positioning along the depth dimension of the user interface container. In some implementations, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending away from the user or the user's viewpoint in different directions and / or from different starting points).In some implementations, when depth is defined relative to the user interface container, the orientation of the depth dimension remains constant relative to the user interface container as the position of the user interface container changes, as the user and / or the user's viewpoint changes (e.g., when multiple different viewers are viewing the same container in a 3D environment, such as during a collaborative session and / or when multiple participants are in a real-time communication session with shared virtual content including the container). In some implementations, for curved containers (e.g., containers including those with curved surfaces or curved content areas), the depth dimension optionally extends into the surface of the curved container. In some cases, z-interval (e.g., the distance between two objects in the depth dimension), z-height (e.g., the distance of one object from another in the depth dimension), z-position (e.g., the position of an object in the depth dimension), z-depth (e.g., the position of an object in the depth dimension), or simulated z-dimensionality (e.g., depth used as a dimension of an object, a dimension of the environment, an orientation in space, and / or an orientation in simulated space) are used to refer to the concept of depth as described above.
[0182] In some implementations, a user may optionally be able to interact with virtual objects in a three-dimensional environment using one or both hands as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system may optionally capture one or both of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above). Alternatively, in some implementations, the user's hands may be seen via the display generating component, through the ability to see the physical environment through the user interface, due to the transparency / semi-transparency of a portion of the user interface being displayed by the display generating component, or due to the projection of the user interface onto a transparent / semi-transparent surface or onto the user's eyes or into the user's field of view. Thus, in some implementations, the user's hands are displayed at corresponding locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if these virtual objects were physical objects in the physical environment. In some implementations, the computer system may update the display of the user's hand representation in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0183] In some embodiments described below, the computer system optionally determines the “effective” distance between a physical object in the physical world and a virtual object in a three-dimensional environment, for example, to determine whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, holding, or within a threshold distance of a virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger pressing a virtual button, a user’s hand grasping a virtual vase, a user’s hand clasped together to pinch / hold the application’s user interface, and two fingers performing any other type of interaction described herein. For example, the computer system optionally determines the distance between a user’s hand and a virtual object when determining whether and / or how a user is interacting with a virtual object. In some embodiments, the computer system determines the distance between a user’s hand and a virtual object by determining the distance between the position of a hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, a user's one or both hands are located at a specific location in the physical world. The computer system optionally captures one or both hands and displays them at a specific corresponding location in a three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if it were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of a virtual object of interest in the three-dimensional environment to determine the distance between the user's one or both hands and the virtual object. In some embodiments, the computer system optionally determines the distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between the user's one or both hands and a virtual object, the computer system optionally determines the corresponding location of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if it were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the user's one or both hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Therefore, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system may optionally perform any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or map the position of the virtual object to the physical environment.
[0184] In some implementations, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what the physical stylus held by the user is pointing at. For example, if the user's gaze is directed at a specific location in the physical environment, the computer system optionally determines a corresponding location in the three-dimensional environment (e.g., a virtual location of the gaze), and if a virtual object is located at that corresponding virtual location, the computer system optionally determines that the user's gaze is directed at that virtual object. Similarly, the computer system may optionally be able to determine the direction in which the physical stylus is pointing in the physical environment based on the orientation of the physical stylus. In some implementations, based on this determination, the computer system determines a corresponding virtual location in the three-dimensional environment corresponding to the location pointed at by the stylus in the physical environment, and optionally determines that the stylus is pointing at the corresponding virtual location in the three-dimensional environment.
[0185] Similarly, the embodiments described herein can refer to the location of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the location of the computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system serves as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to a corresponding location in the three-dimensional environment. For example, the location of the computer system would be its location in the physical environment (and its corresponding location in the three-dimensional environment) such that, if the user stands at that location facing the corresponding portion of the physical environment visible via the display generation component, the user will see from that location objects in the physical environment that are positioned, oriented, and / or sized (e.g., in an absolute sense and / or relative to each other) in the same way as objects displayed or visible in the three-dimensional environment by or via the display generation component of the computer system. Similarly, if the virtual objects displayed in a 3D environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same location as these virtual objects in the 3D environment, and physical objects in the physical environment having the same size and orientation as in the 3D environment), then the position of the computer system and / or the user is the position from which the user will see these virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the 3D environment by the display generation components of the computer system.
[0186] In this disclosure, various input methods are described in relation to interaction with a computer system. When an example is provided using one input device or method, and another example is provided using another input device or method, it should be understood that each example is compatible with and optionally utilizes the input device or method described with respect to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When an example is provided using one output device or method, and another example is provided using another output device or method, it should be understood that each example is compatible with and optionally utilizes the output device or method described with respect to the other example. Similarly, various methods are described in relation to interaction with a virtual or mixed reality environment via a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described with respect to the other example. Therefore, this disclosure discloses embodiments that are combinations of features of a plurality of examples without exhaustively listing all features of the embodiments in the description of each example embodiment.
[0187] User interface and related processes
[0188] Now turn attention to implementations of user interfaces (“UIs”) and associated processes that can be implemented on computer systems (such as portable multifunction devices or head-mounted devices) having display generation components, one or more input devices, and (optionally) one or more cameras.
[0189] Figures 7A to 7J Examples of computer systems that embody a three-dimensional virtual presentation and rehearsal environment associated with a demonstration application, based on some implementation schemes, are illustrated.
[0190] Figure 7A An example is illustrated where a computer system (e.g., an electronic device) 101 displays a three-dimensional environment 702 from the user's viewpoint (e.g., facing the rear wall of the physical environment in which the computer system 101 is located) via a display generation component (e.g., display generation component 120 of FIG. 1). In some embodiments, the computer system 101 includes a display generation component (e.g., a touchscreen) and multiple image sensors (e.g., ...). Figure 3Image sensor 314). The image sensor optionally includes one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a portion of the user (e.g., one or both of the user's hands) when the user interacts with the computer system 101. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display including display generating components for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or sensors for detecting the user's attention (e.g., gaze) (e.g., internal sensors facing inward toward the user's face).
[0191] like Figure 7A As shown, computer system 101 displays a virtual presentation 704A in a three-dimensional environment 702 (e.g., corresponding to 704B in a top view of the three-dimensional environment 702). In some embodiments, virtual presentation 704A includes multiple virtual slides, wherein one or more slides of the virtual presentation include one or more visual content items (e.g., text, photographs, etc.) displayed within the three-dimensional environment 702. In some embodiments, the slides of the virtual presentation are displayed one at a time in the three-dimensional environment, wherein the user of the computer system controls which slide of the virtual presentation is displayed in the three-dimensional environment at any given time (described in further detail below).
[0192] In some implementations, and to facilitate user control over the virtual presentation, computer system 101 displays a speaker notes user interface 706A (e.g., corresponding to 706B in a top-down view of the 3D environment 702). The speaker notes user interface 706A includes one or more optional options for controlling the virtual presentation (described further below). In some implementations, such as Figure 7A As shown in the top view, the speaker notes user interface 706B is displayed in front of and to the sides of the virtual presentation 704B to allow the user to view both virtual presentations simultaneously from their current viewpoints. In some embodiments, the computer system 101 facilitates user control over the slides displayed on the virtual presentation. In some embodiments, the computer system 101 facilitates user control over the slides of the virtual presentation via the speaker notes user interface 706A. For example, upon detecting that a user's gaze is directed towards the speaker notes user interface 706A, and in response to detecting that the user's hand 703A performs an air gesture input, the computer system changes the slides displayed on the virtual presentation 704A, such as... Figure 7BAs shown. In some embodiments, air gesture input includes computer system 101 detecting that hand 703A moves in a specific direction (e.g., to the right or left) when hand 703A performs / holds an air pinch gesture. Based on the detected direction, the computer system advances (e.g., displays the next slide in the slide progression of the virtual presentation) or rewinds (e.g., displays the previous slide in the slide progression of the virtual presentation) the slides being displayed on virtual presentation 704A. Additionally or alternatively, the device advances or rewinds the slides of the virtual presentation in response to detecting that a user's gaze is directed at virtual presentation 708A and that the user's hand 703A performs an air gesture input as described above regarding using the speaker notes user interface 706A to advance and rewind slides.
[0193] In some implementations, the user can move within their physical environment, corresponding to a request to move their viewpoint in a three-dimensional environment, and in response to detecting that the user has moved, the computer system modifies the user's viewpoint, such as... Figure 7C exemplified. like Figure 7C As shown, in response to detecting that user 708 provides input to change their viewpoint of the three-dimensional environment 702, such as by turning their body or head to face a different direction, the computer system modifies the viewpoint of the three-dimensional environment 702 to correspond to the user's detected movement. For example, in Figure 7C In the example, when user 708 changes their viewpoint to look to the left of virtual presentation 704A, the computer system modifies the displayed viewpoint of virtual environment 702 so that virtual presentation 704A appears to be displayed to the right of the user's viewpoint. In some implementations, speaker notes 706A are also displayed to the right of the user's viewpoint, proportional to the change in the user's viewpoint. In some implementations, virtual presentation 704A and speaker notes 706A are world-locked. Additionally or alternatively, and as... Figure 7C As shown, speaker note 706A moves with the user's viewpoint (e.g., speaker note 706A is viewpoint locked).
[0194] Return to Figure 7A For example, in some embodiments, computer system 101 facilitates the display of one or more rehearsal virtual environments (e.g., described in more detail with reference to method 800). A rehearsal virtual environment refers to a three-dimensional environment displayed by the computer system that mimics a real-world physical demonstration environment. In some embodiments, the computer system facilitates a user selecting a rehearsal virtual environment via a speaker notes user interface 706A. For example, the speaker notes user interface 706A includes one or more optional options 724 for selecting a rehearsal virtual environment. In some embodiments, in response to detecting an interest in... Figure 7AIn the optional option 724, select 705. The computer system 101 displays as follows: Figure 7D The rehearsal virtual environment selection user interface 710 is shown. In some implementations, the computer system detects the selection of the optional option 724 by detecting an air pinching gesture made by the user's hand 703A when the user's attention is directed at the optional option 724.
[0195] like Figure 7D As illustrated, in response to detecting a selection of optional option 724, computer system 101 displays a rehearsal virtual environment selection user interface 710 in a three-dimensional environment 702. In some embodiments, the rehearsal virtual environment selection user interface 710 includes one or more optional options for selecting a rehearsal virtual environment from one or more predetermined rehearsal virtual environments stored on the computer system. One or more predetermined rehearsal virtual environments optionally correspond to representations of real-world physical environments. For example, predetermined rehearsal virtual environments include, but are not limited to, conference rooms, lecture halls, and / or auditoriums. In some embodiments, the rehearsal virtual environment selection user interface 710 includes visual indicators associated with predetermined rehearsal virtual environments that can be selected on computer system 101. For example, visual indicators include images associated with predetermined rehearsal virtual environments (e.g., a picture of a conference room associated with a conference room rehearsal virtual environment). One or more visual indicators associated with a predetermined rehearsal virtual environment included in the rehearsal virtual environment selection user interface 710 are optionally selectable, such that in response to detecting one of the selectable visual indicators, the computer system 101 displays the rehearsal virtual environment associated with the selected option. For example, in response to detecting a selection 711 of a visual indicator associated with an auditorium rehearsal virtual environment, the computer system 101 displays the auditorium rehearsal virtual environment, such as... Figure 7E As illustrated. In some implementations, computer system 101 detects the selection of an optional option 711 by detecting the user's hand 703D when the user 708's attention is directed at the optional option.
[0196] Figure 7E An exemplary virtual environment for auditorium rehearsals is illustrated. In some embodiments, computer system 101 displays the virtual auditorium rehearsal environment within a three-dimensional environment 702. The virtual auditorium rehearsal environment includes a visual appearance that mimics a real-world auditorium in which real-world presentations can be demonstrated. For example, the virtual auditorium rehearsal environment includes, for instance, […]. Figure 7E The seating and stairs are illustrated. In some implementations, the computer system 101 initially displays the rehearsal virtual environment from a "presenter's viewpoint," such as... Figure 7EThe auditorium rehearsal virtual environment. The presenter's viewpoint refers to the viewpoint a presenter standing at the front of the auditorium and facing the audience would see in a real-world physical auditorium. Therefore, in some embodiments, and as illustrated in the top view, the auditorium rehearsal virtual environment is displayed such that, from the user 708's viewpoint, the virtual presenter 704B is behind the user 708, while the speaker annotation user interface 706B is in front of the user. In some embodiments, the auditorium rehearsal environment is displayed such that the user's viewpoint faces the audience and speaker annotation 706B, while the virtual presenter 704B is behind the user's viewpoint. In some embodiments, in response to detecting that the user has turned their head in physical space, the computer system 101 updates the orientation of the user's viewpoint, thereby displaying different perspectives of the auditorium rehearsal virtual environment.
[0197] In some implementations, computer system 101 facilitates a change in the viewpoint from which it displays the rehearsal virtual environment. For example, in addition to displaying the rehearsal virtual environment from the presenter's viewpoint as described above, computer system 101 may also display the rehearsal virtual environment from alternative viewpoints, including but not limited to an "audience viewpoint." In some implementations, an audience viewpoint refers to the viewpoint that audience members seated in an auditorium and facing the virtual presentation would see in a real-world physical auditorium. In some implementations, computer system 101 facilitates a change in the viewpoint of the rehearsal environment by displaying one or more optional options for changing the viewpoint of the rehearsal virtual environment. For example, in response to detecting an... Figure 7E The speaker notes at 706A include optional option 726, selection 715, as displayed in computer system 101. Figure 7F The rehearsal virtual environment setup user interface 714 is shown.
[0198] In some implementations, the virtual environment settings user interface 714 includes one or more optional options for adjusting one or more settings associated with the rehearsal virtual environment. For example, the rehearsal virtual environment settings user interface includes optional options for changing viewpoint, lighting settings, and audio settings associated with the rehearsal virtual environment. In some implementations, lighting settings refer to the brightness of the rehearsal virtual environment. In some implementations, audio settings refer to how audio is presented in the virtual environment (e.g., volume, treble, bass, etc.). In some implementations, in response to detecting a selection 717 of the virtual environment settings user interface 714, the computer system 101 displays the rehearsal virtual environment from different viewpoints, such as... Figure 7G As illustrated. In some implementations, computer system 101 detects the selection 717 of an optional option by detecting the user's hand 703F when the user 708's attention is directed at the optional option.
[0199] Figure 7GAn example is shown of a virtual rehearsal environment in an auditorium (described above) displayed from the audience's viewpoint. In some implementations, and as... Figure 7F As illustrated in the top-down view, when the rehearsal virtual environment is displayed from the audience's viewpoint, both the virtual presentation 704B and the speaker notes 706B are displayed in front of the user 708's viewpoint. In some implementations, the user can modify the viewpoint by changing the positioning, such as relative to... Figure 7B As described. Therefore, optionally, user 708 can modify the viewpoint to view the virtual environment from the perspective of a specific audience positioning (e.g., seat) within the virtual environment.
[0200] Return to reference Figure 7D For example, in response to detecting a selection 713 of the optional option in the user interface 710 for the rehearsal virtual environment associated with the meeting room rehearsal virtual environment, the computer system 101 displays as shown in the image. Figure 7H The illustrated virtual environment for meeting room rehearsals. In some implementations, the virtual environment for meeting room rehearsals is configured to have the visual appearance of a real-world meeting room. For example, as... Figure 7H As illustrated, the virtual environment for meeting rehearsals includes a meeting table and chairs displayed in front of the user's viewpoint. In some implementations, the viewpoint associated with the virtual environment for meeting room rehearsals includes a virtual presentation 704A and a speaker notes user interface 706A displayed in front of the user's viewpoint, such as... Figure 7H The top view is further illustrated. Similar to the example of the auditorium rehearsal environment, the viewpoint of the conference room rehearsal virtual environment can also be changed / modified according to the methods described above. In some implementations, in response to detecting that a user has turned their head in the physical space, the computer system 101 updates the orientation of the user's viewpoint, thereby displaying different perspectives of the auditorium rehearsal virtual environment. Compared to the auditorium rehearsal virtual environment, the conference room rehearsal virtual environment is smaller in space, and therefore, the virtual presentation 704A appears closer to the user 708's viewpoint than it would be in the audience's viewpoint in the auditorium rehearsal virtual environment.
[0201] In some implementations, in addition to facilitating navigation of the virtual presentation and rehearsal environment as described above, the speaker notes user interface 706A may also include one or more optional tools for interacting with the virtual presentation in the rehearsal virtual environment. For example, as Figure 7H As illustrated, the speaker notes user interface 706A includes an optional option 728 for enabling a virtual laser pointer. In some embodiments, and in response to detecting a selection 719 of the optional option 728, the computer system 101 displays a virtual laser pointer that can be moved by the user, such as... Figure 7IAs illustrated. In some implementations, the computer system 101 detects a selection 719 of the optional option 728 by detecting an air pinching gesture made by the user's hand 703F when the user's attention is directed at the optional option 728.
[0202] like Figure 7I As illustrated, computer system 101 displays a virtual laser pointer 722 in response to selection 719 of optional option 728. In some embodiments, when the virtual laser pointer is activated, computer system 101 displays a visible line from the user's viewpoint (and / or as if emanating from the user's hand or fingers) to a portion of the virtual presentation. Additionally or alternatively, computer system 101 displays a point on the presentation where the line would hit the virtual presentation. In some embodiments, when computer system 101 detects movement of a part of the user's body (such as movement of the user's hand), it modifies the positioning of the visible line and / or point within the rehearsed virtual environment. In some embodiments, the positioning / or orientation of the line and / or point corresponds to the positioning and / or orientation of the user's hand. For example, and as... Figure 7J As illustrated, if the user's hand 703J is detected to be moving to the left, the positioning of the visible line will also be moved to the left proportionally to the detected movement of the user's hand. Additionally or alternatively, the visible point is displayed on the virtual presentation at a location based on the positioning of a part of the user's body, thus visually indicating that the virtual laser pointer is pointing to a specific part of the virtual presentation. In some embodiments, the line is generated as if it extends through / from the user's forearm and from the palm or other part of the user's hand. In some embodiments, and when the user activates the virtual laser pointer, the computer system 101 responds to detecting air pinch and drag gestures when the user's attention is directed to the virtual presentation without navigating through the slides of the virtual presentation. Optionally, such air gestures pointing to the speaker's notes 706A will cause the slides of the virtual presentation 704A to move forward or backward according to the detected movement of the user's hand.
[0203] Figure 8 This is a flowchart illustrating a process for displaying a virtual environment associated with a demonstration application, according to some embodiments. In some embodiments, method 800 is performed at a computer system (e.g., computer system 101 in Figure 1, such as a tablet, smartphone, wearable computer, or head-mounted device), which includes display generation components (e.g., Figure 1, ...). Figure 3 and Figure 4The display generating component 120 (e.g., a heads-up display, a monitor, a touchscreen, and / or a projector) and one or more cameras (e.g., a camera pointing downwards at the user's hand (e.g., a color sensor, an infrared sensor, or other depth-sensing camera) or a camera pointing forward from the user's head). In some embodiments, method 800 is performed by storing in a non-transitory computer-readable storage medium and by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., ...). Figure 1A The control unit 110 in the middle executes instructions to manage. Some operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0204] In some embodiments, method 800 is performed at a computer system communicating with a display generating component and one or more input devices. These include, for example, mobile devices (e.g., tablets, smartphones, media players, or wearable devices) or computers or other electronic devices. In some embodiments, the display generating component is a display integrated with an electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users. In some embodiments, the one or more input devices include the ability to receive user input (e.g., capture user input, detect user input, etc.) and transmit information associated with that user input to the computer system. Examples of input devices include touchscreens, mice (e.g., external), touchpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the computer system), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye-tracking devices, and / or motion sensors (e.g., hand-tracking devices, hand motion sensors), etc. In some implementations, the computer system communicates with a hand-tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreens, touchpads)). In some implementations, the hand-tracking device is a wearable device, such as a smart glove. In some implementations, the hand-tracking device is a handheld input device, such as a remote control or stylus.
[0205] In some implementations, when a virtual presentation associated with a demonstration application (such as...) is displayed in a first three-dimensional environment via a display generation component... Figures 7A to 7JWhen a virtual presentation (704A) is being presented, the computer system receives (802a) a first input (e.g., a pinch-to-open gesture, gaze, or tap on a touchscreen) via one or more input devices corresponding to a request to display the virtual presentation in a second mode different from the first mode of the presentation application. Figure 7D In this embodiment, the first three-dimensional environment includes a portion of the physical environment of the user of the computer system, and the virtual presentation is presented in a first mode of the presentation application. In some embodiments, the first three-dimensional environment incorporates at least a partial representation of the user's real-world physical environment when using the computer system (e.g., via active or passive pass-through). In some embodiments, presenting the virtual presentation in the first mode of the presentation application includes presenting the virtual presentation in a three-dimensional environment that includes a portion of the user's physical environment. Optionally, when in the first mode, the virtual presentation can be edited by the user. Additionally or alternatively, when in the first mode, the virtual presentation is displayed as a slideshow and cannot be edited by the user. In some embodiments, a first input corresponding to a request to display the virtual presentation in a second mode is detected by the device via optional options presented in a user interface displayed in the first three-dimensional environment when the presentation application is in the first mode. In some embodiments, the second mode includes displaying the virtual presentation in a rehearsal virtual environment (described in further detail below). Optionally, the rehearsal environment is configured to mimic a real-world physical environment in the virtual three-dimensional environment.
[0206] In some implementations, in response to receiving the first input (802b), the computer system initiates (802c) a process of displaying a virtual presentation in a corresponding (optionally rehearsed) virtual environment different from the first three-dimensional environment, such as in Figure 7EIn this embodiment, when the virtual presentation is displayed in a corresponding (optionally rehearsed) virtual environment, that portion of the user's physical environment is invisible via the display generation components. In some embodiments, the three-dimensional environment is an extended reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment, and the virtual presentation is displayed within the three-dimensional environment. In some embodiments, displaying the virtual presentation in a first presentation mode includes displaying the virtual presentation in a manner less than full immersion. Immersion level includes the degree to which virtual content displayed by a computer system occludes background content around / behind the virtual content (e.g., a 3D environment including the physical environment), the background content optionally being a representation of the virtual content and / or the user's physical environment, optionally including the number of items of displayed background content and the visual characteristics (e.g., color, contrast, and / or opacity) of the displayed background content, and / or the angular range of the virtual content displayed via display generation components (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, and / or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view consumed by the virtual content displayed via display generation (e.g., 33% of the field of view consumed by the virtual content at low immersion, 66% of the field of view consumed by the virtual environment at medium immersion, and / or 100% of the field of view consumed by the virtual content at high immersion). In some embodiments, at a first (e.g., high) immersion level, the background, virtual and / or real objects are displayed in an occluded manner. For example, corresponding virtual content with a high level of immersion is displayed without concurrently displaying background content (e.g., in full-screen or fully immersive mode). In some embodiments, at a second (e.g., lower) level of immersion, the background, virtual, and / or real objects are displayed in an unobstructed manner (e.g., not dimmed, not blurred, and / or not removed from the display). For example, virtual content with a low level of immersion is optionally displayed concurrently with background content, which is optionally displayed at full brightness, color, and / or semi-transparency. As another example, virtual content displayed at a medium level of immersion is optionally displayed concurrently with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects differ between background objects. For example, at a particular level of immersion, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are not displayed. In some embodiments, the rehearsed virtual environment is displayed with full immersion. In some embodiments, the rehearsed virtual environment includes a virtual 3D representation of the physical environment in which a virtual demonstration can be performed. For example, sample rehearsal virtual environments include virtual meeting rooms, virtual auditoriums, and virtual classrooms.In some implementations, the computer system displays a user interface including multiple optional options, each corresponding to a type of rehearsal virtual environment. When the device detects that a user has selected one of the optional options, the device optionally displays the rehearsal virtual environment associated with the user-selected option. Additionally or alternatively, in some implementations, the process of initiating the display of a virtual presentation in the corresponding rehearsal virtual environment includes: stopping the display of a first three-dimensional environment; and displaying a second three-dimensional environment (e.g., the rehearsal virtual environment) associated with the user-selected option as detected at the computer system. Optionally, stopping the display of the first three-dimensional environment includes stopping the display of that portion of the user's physical environment. For example, in response to a first input, the computer system stops displaying (actively or passively) the user's physical environment and initiates a process of displaying a fully virtual three-dimensional environment based on the detected input. In some implementations, displaying the rehearsal virtual environment causes the virtual presentation to appear behind the user's viewpoint (e.g., outside the user's field of view). Optionally, displaying the rehearsal virtual environment causes the virtual presentation to appear in front of the user's viewpoint (e.g., within the user's field of view). Optionally, in the second mode, the device allows the presentation to be edited by the user. Displaying a rehearsal virtual environment when selected by the user allows the device to minimize the amount of time required to render a fully virtual environment, thereby saving computational resources of the computing system.
[0207] In some implementations, in response to receiving a first input, the computer system displays a virtual environment selection user interface via a display generation component for selecting one or more pre-defined (optionally rehearsed) virtual environments for a demo application, such as... Figure 7D Interface 710. In some embodiments, the rehearsal virtual environment selection user interface includes images and / or other identifying information associated with each of the one or more predefined rehearsal virtual environments that can be selected. Examples of predefined rehearsal virtual environments include, but are not limited to, auditoriums, conference rooms, and lecture halls.
[0208] In some implementations, when the virtual environment selection user interface is displayed, the computer system receives, via one or more input devices, a second input corresponding to the selection of one or more predetermined virtual environments, such as... Figure 7D User input 711. In some implementations, the device detects the selection of identifying information (e.g., a tap or air gesture / pinch) associated with a predefined rehearsal virtual environment.
[0209] In some implementations, in response to receiving a second input, the computer system displays a virtual presentation in one or more (optionally rehearsed) virtual environments, such as in Figure 7E In some implementations, once the device has detected the selection of one of the optional options for the rehearsal virtual environment, the device stops displaying the current 3D environment in which the virtual demo is displayed and instead displays the 3D environment associated with the selected predefined rehearsal virtual environment. In some examples, the virtual demo is presented in edit mode when it is first displayed in the selected rehearsal virtual environment, where the user is able to edit the content of the demo. Alternatively, the virtual demo is displayed in demo mode, where the user cannot edit the content of the virtual demo. Providing predefined rehearsal virtual environments reduces the need for users to create custom virtual environments in which virtual demos are presented, thereby saving both the computing and memory resources associated with creating custom rehearsal virtual environments.
[0210] In some implementations, displaying a virtual presentation in a selected, predetermined (optionally rehearsed) virtual environment by a computer system includes displaying the selected, predetermined (optionally rehearsed) virtual environment from a viewpoint corresponding to the presenter's viewpoint within the selected, predetermined (optionally rehearsed) virtual environment, such as in... Figure 7E In some implementations, the presenter's viewpoint involves displaying the rehearsal environment from the perspective of the presenter of the virtual presentation, such that the virtual audience of the virtual presentation is in front of the user, while the virtual presentation is placed behind the user's viewpoint, thereby simulating a real-world scenario where the presenter is demonstrating a presentation to the audience. In some implementations, when initially in the presenter's viewpoint, the device detects movement of the user's viewpoint / perspective within the rehearsal virtual environment and adjusts the viewpoint of the virtual environment accordingly. Therefore, if the rehearsal virtual environment is displayed at the presenter's viewpoint and the computing system detects that the user has turned or moved within their physical space such that their viewpoint has changed to face the virtual presentation, the computing system will display the virtual presentation in front of the user, and the audience will be behind the user. Initially displaying the presentation in the presenter's viewpoint when the rehearsal virtual environment is initially displayed reduces the likelihood that the user will need to change their viewpoint once the rehearsal virtual environment is displayed, thereby saving both computing and memory resources of the computer system.
[0211] In some implementations, displaying a virtual presentation in response to a second input within a selected, predetermined (optionally rehearsed) virtual environment includes: displaying a virtual environment settings user interface via a display generation component for changing one or more settings associated with the selected, predetermined (optionally rehearsed) virtual environment, such as... Figure 7FThe user interface 714 is included. In some embodiments, the virtual environment settings user interface includes one or more optional options for changing one or more settings associated with the rehearsal environment. For example, when the device detects a selection of an optional option included in the virtual environment settings user interface, the device modifies an aspect of the rehearsal virtual environment associated with the optional option. In some embodiments, the settings user interface includes optional options for changing the appearance of the rehearsal virtual environment, the audio presented in the rehearsal virtual environment, and the viewpoint of the rehearsal virtual environment. In some embodiments, the settings user interface is automatically displayed when the device initially displays the rehearsal virtual environment. Additionally or alternatively, the settings user interface is displayed in response to the device detecting input from a user. Displaying the settings user interface in the rehearsal virtual environment minimizes the amount of user input required to modify the rehearsal environment, thereby saving computational resources associated with modifications to the rehearsal environment.
[0212] In some implementations, when a virtual environment settings user interface is displayed, the computer system receives third input directed to the virtual environment settings user interface via one or more input devices to change one or more virtual lighting settings associated with a selected, predetermined (optionally rehearsed) virtual environment. In some implementations, in response to receiving the third input for changing one or more virtual lighting settings, the computer system adjusts the virtual lighting characteristics of the selected, predetermined (optionally rehearsed) virtual environment according to the third input, such as when the device detects an error in the virtual lighting settings. Figure 7F In the case of selecting the "Visual Options" optional option as part of the user interface 714, as illustrated, in some embodiments, the virtual lighting settings include the brightness (e.g., virtual light characteristics) of the rehearsal virtual environment associated with the amount of illumination present in the rehearsal virtual environment. In some embodiments, the device detects modifications to the lighting settings of the rehearsal virtual environment and modifies the brightness of the virtual environment based on the detected modifications. For example, the device modifies the overall brightness of the rehearsal virtual environment by adjusting the intensity of one or more virtual lights in the rehearsal virtual environment. In some embodiments, the position of one or more virtual lights is based on the real-world physical environment associated with the rehearsal virtual environment. Including the lighting settings as part of the virtual environment settings user interface minimizes the amount of user input required to modify the lighting settings of the rehearsal virtual environment, thereby saving computational resources associated with modifications to the rehearsal environment.
[0213] In some implementations, when the virtual environment settings user interface is displayed, the computer system receives, via one or more input devices, a third input corresponding to a request to change the viewpoint of the selected, pre-defined (optionally rehearsed) virtual environment to the presenter's viewpoint, such as in Figure 7FIn some implementations, in response to receiving a third input, and based on determining that a virtual environment can be displayed (optionally rehearsed) from a viewpoint different from the presenter's viewpoint, the computer system displays a selected predetermined (optionally rehearsed) virtual environment from the presenter's viewpoint, wherein displaying the selected predetermined (optionally rehearsed) virtual environment from the presenter's viewpoint includes: placing the virtual presentation behind the user's viewpoint within the selected (optionally rehearsed) virtual environment, such as in... Figure 7E In some implementations, the virtual environment settings user interface includes optional options for displaying the rehearsal virtual environment from the presenter's viewpoint (described in detail above). In some implementations, and when the rehearsal virtual environment is not yet in the presenter's viewpoint (described in detail above), the device detects the selection of an optional option associated with displaying the rehearsal virtual environment in the presenter's viewpoint and modifies the viewpoint of the rehearsal virtual environment according to the presenter's viewpoint described above. Including optional options on the virtual environment settings interface to display the rehearsal virtual environment in the presenter's viewpoint minimizes the amount of user input required to modify the viewpoint of the rehearsal environment, thereby saving both computational and memory resources of the computer system.
[0214] In some implementations, when the virtual environment setup user interface is displayed, the computer system receives a fourth input via one or more input devices corresponding to a request to change the viewpoint of the selected, pre-defined (optionally rehearsed) virtual environment to the viewpoint of the audience, such as an input applied to... Figure 7F User input 717 in the user interface 714. In some embodiments, in response to receiving a fourth input, and based on determining that a virtual environment can be displayed (optionally rehearsed) from a viewpoint different from the viewpoint of the audience, the computer system displays a predetermined (optionally rehearsed) virtual environment selected from the viewpoint of the audience, wherein displaying the selected predetermined (optionally rehearsed) virtual environment from the viewpoint of the audience includes: displaying a virtual presentation in the selected (optionally rehearsed) virtual environment in front of the user's viewpoint. In some embodiments, the virtual environment settings user interface includes an optional option to display the rehearsed virtual environment from the viewpoint of the audience. In some embodiments, the viewpoint of the audience includes displaying the virtual presentation (from the user's viewpoint) at a position in front of the user so that an audience member simulating the virtual presentation would view the presentation while it is being presented. In some embodiments, and if the rehearsed virtual environment is not yet in the viewpoint of the audience, the device detects the selection of an optional option associated with displaying the rehearsed virtual environment in the viewpoint of the audience, and modifies the viewpoint of the rehearsed virtual environment according to the viewpoint of the audience described above. The virtual environment settings interface includes optional options to display the rehearsal virtual environment from the audience's viewpoint, minimizing the amount of user input required to modify the viewpoint of the rehearsal environment, thereby saving both computing and memory resources of the computer system.
[0215] In some implementations, the virtual environment settings user interface includes one or more interactive options that are interactive in response to detecting input from a first part of the user when the user's attention is directed to the optional option, such as in Figure 7F The illustrated user interface 714 is in an interactive configuration. In some implementations, the virtual environment setup user interface can be navigated using the user's gaze and / or one or more air gestures such as air pinch (described above). For example, to select an optional option on the setup user interface, the device can detect that the user is performing an air pinch while gazing at the optional option. The computing system running the demonstration application optionally includes components capable of detecting the user's real-world movement and gaze (described above), and when the computing system determines that the user is interacting with the virtual environment setup user interface, the system performs actions commensurate with the detected interaction. Detecting input from a portion of the user when the user's attention is directed at an optional option minimizes the number of input devices required to receive input from the user and minimizes the possibility of erroneous user input, thereby saving computing resources associated with managing input devices and correcting erroneous input.
[0216] In some implementations, when the virtual settings user interface is displayed in an expanded state, the computer system receives a third input directed to the virtual settings user interface via one or more input devices, wherein the third input includes input from the user's first portion directed to the first portion of the virtual settings user interface, followed by movement of the user's first portion in a downward direction, such as in... Figure 7F The user interface 714 is interactive and the user minimizes or collapses the user interface. In some embodiments, the virtual setup user face is interactive to expand and minimize the user interface. In some embodiments, the interaction includes detecting when the user is looking at a visible portion of the setup user interface to expand and minimize the interface. For example, if the virtual setup user interface is minimized and the computing system detects the user looking and pinching and pulling up the user interface, the computing system expands the user setup interface to display one or more selectable / interactive options associated with the interface. Optionally, if the virtual setup user interface is expanded and the computing system detects the user pinching and pulling down the user interface, the computing system minimizes the user setup interface, thereby stopping the display of at least some or all of the selectable / interactive options associated with the interface. Allowing the user to expand and minimize the virtual environment setup user interface ensures that parts of the rehearsal virtual environment are not obscured or covered by the interface when the user attempts to interact with those parts, thereby minimizing erroneous user input and saving computing system resources associated with correcting erroneous user input.
[0217] In some implementations, when displaying the virtual environment settings user interface, the computer system receives third input (such as...) via one or more input devices. Figure 7H Input 728 on the speaker annotation interface 706A in the system is used to activate a virtual laser pointer in the selected (optionally rehearsed) virtual environment. In some embodiments, in response to receiving a third input for activating the virtual laser pointer, the computer system displays a virtual laser point via a display generation component, such as in... Figure 7I In this context, the laser pointer moves within a selected (optionally rehearsed) virtual environment based on the user's detected movement in the first part, such as in... Figure 7J In some embodiments, the virtual laser, when activated, displays a visible line from the user's viewpoint to a portion of the virtual presentation. Alternatively, the device does not display the visible line and instead displays a dot on the presentation where the line would hit the virtual presentation. In some embodiments, the positioning of the visible line is modified when the device detects movement of a part of the user's body (such as movement of the user's hand). For example, if the user's hand is detected to be moving to the left, the positioning of the visible line will also be shifted to the left proportionally to the detected movement of the user's hand. Alternatively, a visible dot is displayed on the virtual presentation at a location based on the positioning of a part of the user's body, thus visually indicating that the virtual laser pointer is pointing to a specific portion of the virtual presentation. In some embodiments, the line is generated as if it extends through / from the user's forearm and from the palm or other part of the user's hand. In some embodiments, one or more visual indicators (e.g., visible lines and visible dots) associated with the virtual laser pointer are displayed based on how the real-world physical laser point would appear when the presentation is made in the real world. In some implementations, when a user activates the virtual laser pointer, the virtual laser can move independently of the user's hand state (e.g., the user's hand can be in a "ready" state or another state or hand shape). In some implementations, when the hand is detected as being pinched by the device, the device displays the virtual laser pointer line with higher visual salience than when the hand is detected as not being pinched. Providing a virtual laser pointer allows the computing device to more closely mimic how a demonstration would be presented in a real-world setting and minimizes the amount of user input required to activate the virtual laser pointer by providing an optional option to activate the virtual laser pointer on the virtual environment setup user interface, thereby minimizing erroneous user input and thus saving computational resources associated with correcting erroneous user input.
[0218] In some implementations, the computer system receives a fourth input via one or more input devices, the fourth input including an air gesture from the user pointing at the virtual presentation. In some implementations, in response to detecting the fourth input, and based on determining that the virtual laser pointer is not activated, the computer system navigates through the virtual presentation based on the fourth input, such as in... Figure 7B In some implementations, upon determining that the virtual laser pointer is activated, the computer system abandons navigation through the virtual presentation based on a fourth input. In some implementations, and when moving the virtual laser pointer in a rehearsed virtual environment by detecting movement of a part of the user's body, the device prevents the user from using air gestures to navigate through the virtual presentation in order to minimize any errors associated with the user intending to move the virtual laser pointer and, alternatively, interpreting the movement as navigation of the content of the virtual presentation. For example, when the virtual laser pointer is not enabled, the user navigates the presentation by looking at the virtual presentation or speaker notes associated with the virtual presentation, performing air gestures such as pinching in the air, and dragging in a given direction. However, when the virtual laser pointer is enabled, the user will not be able to use the gestures described above to navigate the presentation (e.g., the computer system will not respond to detecting such gestures when the laser pointer is enabled and navigate through the virtual presentation). In some implementations, when the virtual laser pointer is not activated, any air gestures detected by the computer system are interpreted as requests to navigate the content of the virtual presentation because the possibility of erroneous representation of the user's movement is minimized. When the virtual laser pointer is activated, preventing users from using air gestures to navigate the content of the virtual presentation minimizes the possibility of the device misinterpreting the air gestures, and thus saves computing resources associated with correcting any misinterpretations of the air gestures.
[0219] In some implementations, displaying a virtual presentation in a corresponding (optionally rehearsed) virtual environment includes: based on determining that the virtual presentation is in presentation mode, displaying a speaker notes user interface via a display generation component for displaying information associated with the virtual presentation, such as... Figure 7AThe presenter notes user interface 706A is described above. In some embodiments, when the device detects that the virtual presentation is in presentation mode, the virtual presentation is displayed in a non-editable format to mimic the real-world scenario in which the presentation will be delivered. In the real-world scenario, presenters may have presenter notes that they can use to provide information about the presentation, such as the content of the slides, as well as other notes that the presenter does not want the audience to see. In some embodiments, the presenter notes user interface includes one or more visual representations of the slides associated with the virtual presentation. To further mimic a real-world presentation, the device displays a presenter notes user interface that displays information similar to that available to a presenter using real-world presenter notes. In some embodiments, and as described in further detail below, the presenter notes user interface includes one or more interactive / optional options that, when selected, allow the device to advance or rewind the slides being displayed on the virtual presentation. Providing a speaker notes user interface that allows users to interact with forward and backward slides on a virtual presentation minimizes erroneous user input associated with not being able to view content in the presentation other than the slides currently being displayed on the virtual presentation, thereby saving computational resources associated with correcting erroneous user input.
[0220] In some implementations, the speaker notes user interface is displayed at a location within the corresponding (optionally rehearsed) virtual environment that differs from the location where the virtual presentation is displayed within the corresponding (optionally rehearsed) virtual environment, such as in Figure 7A In some implementations, the speaker notes are placed in a location different from where the virtual presentation is displayed in the rehearsal virtual environment to prevent the speaker notes user interface from obscuring any part of the virtual presentation. By placing the speaker notes user interface in a location different from where the virtual presentation is displayed in the rehearsal virtual environment, users can view the speaker notes and the virtual presentation simultaneously in the same way as in a real-world presentation context. In some implementations, the speaker notes user interface is displayed to the side, below, or above the virtual presentation, allowing users to view both the virtual presentation and the speaker notes user interface simultaneously without one obscuring the other. Additionally or alternatively, the speaker notes user interface is displayed in a location that only partially obscures the virtual presentation, allowing users to view both the speaker notes and most of the virtual presentation simultaneously. Displaying the speaker notes user interface in a location different from where the virtual presentation is displayed in the rehearsal virtual environment minimizes erroneous user input associated with speaker notes that obscure the user's view of the virtual presentation while viewing the speaker notes user interface, and thus saves memory and computational resources associated with correcting erroneous user input.
[0221] In some implementations, when the speaker notes user interface is displayed, the computer system receives a second input from a first portion of the user via one or more input devices. This second input includes a first air pinch pointing towards a portion of the speaker notes user interface, followed by movement of the user's first portion. In some implementations, in response to receiving the second input, the computer system moves the speaker notes user interface within a corresponding (optionally rehearsed) virtual environment based on detected movement of the user's first portion, such as when the computer system detects an air pinch when the user's attention is directed towards speaker notes 706A. Figure 7A In the case of speaker notes 706A, and subsequently, the user's hand is moved. In some embodiments, although the speaker notes user interface is initially displayed in a different location within the rehearsal virtual environment than the virtual presentation, allowing the user to view both the speaker notes user interface and the virtual presentation simultaneously, the speaker notes can move within the rehearsal virtual environment. In some embodiments, the speaker notes user interface is moved when the device detects that the user is looking at the speaker notes user interface and performs an air pinch pointing to a part of the speaker notes user interface and detects subsequent movement of a part of the user's body (such as the user's hand). The device moves the position of the speaker notes user interface based on the movement of that part of the user's body (e.g., in the direction corresponding to the movement of that part of the user's body and / or with a magnitude corresponding to the magnitude of the movement of that part of the user's body). Allowing the speaker notes user interface to move within the rehearsal virtual environment allows for user customization of the rehearsal virtual environment with minimal user input and minimal erroneous user input, thus saving computational resources associated with customization of the rehearsal virtual environment. In some implementations, when displaying a virtual presentation, the computer system receives a second input from a portion of the user via one or more input devices. This second input includes a first air gesture from a first portion of the user, the first air gesture including movement of the first portion in a corresponding direction. In some implementations, in response to receiving the second input, based on determining that the corresponding direction is a first direction, the computer system navigates through the virtual presentation in the first corresponding direction, such as in… Figure 7B middle.
[0222] In some implementations, based on the determination that the corresponding direction is a second direction different from the first corresponding direction, the computer system navigates through the virtual presentation in a second corresponding direction different from the first corresponding direction, such as when a user moves their hand in the opposite direction to move the slide backward instead of forward. Figure 7AIn scenarios where slides are moving forward or backward, the device advances or rewinds slides based on detected air gestures from a part of the user's body, such as the user's hand. In some implementations, if the device detects a user's gaze pointing at the virtual presentation, and while the user is looking at the presentation, it detects an air pinch from the user's hand, and if it detects that the user's hand (while maintaining the air pinch) moves from left to right, the device advances the slide currently displayed on the virtual presentation (e.g., the next slide in a series of slides associated with the virtual presentation). In some implementations, the slides do not advance solely due to detected movement of the user's hand. In some implementations, if the device detects that the user's hand moves from right to left (e.g., in the opposite direction), it rewinds the slide currently displayed on the virtual presentation (e.g., the previous slide in a series of slides associated with the virtual presentation). Using detected air gestures to navigate content displayed on a virtual presentation in a rehearsed virtual environment minimizes the number of input devices required to interact with the virtual presentation, thereby saving computational resources associated with accepting input from multiple input devices.
[0223] In some implementations, the first input refers to a speaker notes user interface displayed concurrently with the virtual presentation in the corresponding (optionally rehearsed) virtual environment, such as in Figure 7A In some implementations, the device advances or rewinds slides displayed on the virtual presentation based on a user's gaze and air gestures (described above) pointing at the speaker notes user interface. As described above, the speaker notes user interface includes one or more visual representations of slides associated with the virtual presentation, and therefore the device navigates the content of the virtual presentation (e.g., advances and rewinds slides) by detecting air gestures pointing at the speaker notes user interface and specifically at the visual representations of slides included on the speaker notes user interface. Detecting air gestures pointing at the speaker notes user interface and displaying slides on the virtual presentation based on the detected air gestures allows the device to minimize the number of input devices required to navigate the virtual presentation, thereby saving computational resources associated with accepting input from multiple input devices.
[0224] In some implementations, the first input points to a virtual presentation in a corresponding (optionally rehearsed) virtual environment, such as when the user will... Figure 7A The input pointer in Figure 7AIn the case of virtual presentation 704A, in some embodiments, the device advances or rewinds the slides displayed on the virtual presentation based on the user's gaze and air gestures (described above) pointing at the virtual presentation. As described above, the device displays slides associated with the virtual presentation. In some embodiments, the device navigates the content of the virtual presentation (e.g., advances and rewinds the slides) by detecting air gestures pointing at the virtual presentation in the rehearsal virtual environment. Detecting air gestures pointing at the virtual presentation and displaying slides on the virtual presentation based on the detected air gestures allows the device to minimize the number of input devices required to navigate the virtual presentation, thereby saving computational resources associated with accepting input from multiple input devices.
[0225] Figures 9A to 9E Examples of computer systems that illustrate a three-dimensional virtual representation of virtual objects associated with a display and demonstration application according to some implementation schemes are provided.
[0226] Figure 9A An example is illustrated where a computer system (e.g., an electronic device) 101 displays a three-dimensional environment 902 from the user's viewpoint (e.g., facing the rear wall of the physical environment in which the computer system 101 is located) via a display generation component (e.g., display generation component 120 of FIG. 1). In some embodiments, the computer system 101 includes a display generation component (e.g., a touchscreen) and multiple image sensors (e.g., ...). Figure 3 Image sensor 314). The image sensor optionally includes one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a portion of the user (e.g., one or both of the user's hands) when the user interacts with the computer system 101. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display including display generating components for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or sensors for detecting the user's attention (e.g., gaze) (e.g., internal sensors facing inward toward the user's face).
[0227] like Figure 9AAs shown, computer system 101 (e.g., as described with respect to methods 800 and 1000) displays a virtual presentation 904A (e.g., corresponding to 904B in a top view of the three-dimensional environment 902) in a three-dimensional environment 902. In some embodiments, the virtual presentation 904A includes a plurality of virtual slides, wherein each slide of the virtual presentation includes one or more visual content items (e.g., text, photographs, etc.) displayed within the three-dimensional environment 902. In some embodiments, the slides of the virtual presentation are displayed one at a time in the three-dimensional environment, wherein the user of the computer system controls which slide of the virtual presentation is displayed in the three-dimensional environment at any given time.
[0228] In some embodiments, and to facilitate user control over the virtual presentation, computer system 101 displays a speaker notes user interface 906A (e.g., corresponding to 906B in a top view of the three-dimensional environment 902). The speaker notes user interface 906A includes one or more optional options for controlling the virtual presentation (described further below). In some embodiments, the virtual presentation 904A is displayed on a three-dimensional or two-dimensional surface within the three-dimensional environment 902, such as... Figure 9A As illustrated, the content displayed on virtual presentation 904A will be shown in two dimensions. However, in some embodiments, virtual presentation 904A includes one or more three-dimensional objects that can be displayed within a three-dimensional environment 902. Since the content of virtual presentation 904A is displayed in two dimensions, virtual presentation 904A may include one or more two-dimensional representations of three-dimensional virtual objects that can be displayed in three dimensions in response to input from a user in one or more user interfaces associated with virtual presentation 904A (described in further detail below). For example, virtual object 908 (e.g., an octopus) displayed as part of virtual presentation 904A is a two-dimensional representation of a three-dimensional object (displayed from the side view of a three-dimensional octopus).
[0229] In some implementations, in response to the detection of certain user input, computer system 101 may use a three-dimensional representation to display virtual object 908 as an alternative and / or supplement to a two-dimensional representation of virtual object 908 displayed on virtual presentation 904A. For example, in response to the detection of a user's hand 903A providing an air pinch gesture when the user's attention (e.g., gaze 921) is directed at virtual object 908, such as... Figure 9A As shown, the computer system displays visual indicators at and / or around the virtual object 908, such as... Figure 9B As illustrated. In some implementations, and as... Figure 9BAs illustrated, the visual indicator 910 is configured to provide a visual representation that the virtual object 908 is a two-dimensional representation of a three-dimensional virtual object. In addition to indicating that the virtual object 908 has been selected, this visual representation can also be displayed using a three-dimensional representation. If the user selects a virtual object that is not a three-dimensional virtual object, the visual indicator may optionally not be displayed for the virtual object. In some embodiments, the visual indicator 910 includes a box surrounding the virtual object, such as... Figure 9B As illustrated. However, the example of visual indicator 910 should not be considered limiting, and other visual indicators (such as color, highlighting, or text) may be applied to a 3D virtual object to indicate that the virtual object is a 3D virtual object. In some embodiments, a user may use visual indicator 910 to manipulate virtual object 908 to change the orientation in which virtual object 908 is displayed on virtual presentation 904A. For example, in some embodiments, in response to the detection of user input (e.g., a pinch gesture in the air when the user's attention (e.g., gaze 921) is directed at virtual object 908, and movement of the user's hand 903B), computer system 101 displays virtual object 908 in a new orientation commensurate with the user's detected movement. For example, a two-dimensional representation of a 3D object may be displayed from a new perspective (shown in two dimensions) based on user input. In some embodiments, and when a two-dimensional representation (e.g., virtual object 908) is selected, the device facilitates the user to perform other manipulations of the display of virtual object 908, such as resizing the object, moving the object in virtual presentation 904A, deleting the virtual object, etc.
[0230] In some embodiments, in addition to displaying a visual indicator 910 at the virtual object 908, and in response to the user input described above (e.g., an air pinch gesture when the user's attention (e.g., gaze 921) is directed at the virtual object 908), the computer system 101 displays a "3D View Button" to initiate the process of displaying a three-dimensional representation of the virtual object 908. In some embodiments, the "3D View Button" is displayed as part of a three-dimensional playback user interface 912 that is displayed in response to the user input described above (e.g., an air pinch gesture when the user's attention (e.g., gaze 921) is directed at the virtual object 908), such as... Figure 9B As illustrated. In some embodiments, the 3D playback user interface includes a “3D view button” as an optional option 914 on interface 912. Additionally or alternatively, computer system 101 displays the “3D view button” on the virtual object. In some embodiments, in response to detecting user input (e.g., an air pinch gesture when the user’s attention (e.g., gaze 921) is directed to optional option 914), computer system 101 displays a 3D representation of the virtual object 908 in the 3D environment 902, such as Figure 9C exemplified.
[0231] In some embodiments, and in response to detecting user input as described above, the computer system displays a three-dimensional representation 908A of the virtual object 916 (e.g., corresponding to 916B in a top view of the three-dimensional environment 902). As shown in the top view of the three-dimensional environment, the three-dimensional representation 916B is displayed in front of the virtual presentation 904B at a location different from the virtual presentation 904B, allowing the user to view both the virtual presentation 904A and the three-dimensional representation 916A simultaneously. In some embodiments, when the virtual object 908 is displayed inside the presentation (e.g., without the three-dimensional representation), the virtual object 908 and the virtual presentation 904A are displayed at the same distance from the user's viewpoint. Optionally, when the three-dimensional representation 916A is displayed, the three-dimensional representation and the virtual presentation 904A are displayed at different distances from the user's viewpoint. In some embodiments, the computer system 101 displays the three-dimensional representation 916A with the same orientation as when the virtual object was displayed when the user initiated the display of the three-dimensional representation. For example, if the two-dimensional representation corresponds to a first perspective view of the virtual object 908, then a three-dimensional representation 916A is displayed such that the view of the object from the user's current viewpoint is the first perspective view. Optionally, if the two-dimensional representation corresponds to a second perspective view of the virtual object 908 (different from the first perspective view), then a three-dimensional representation 916A is displayed such that the view of the object from the user's current viewpoint is the second perspective view. In some embodiments, the three-dimensional representation may move around the three-dimensional environment; however, when the three-dimensional view terminates (as described below), the virtual object 908 reappears in its original position in the virtual presentation 904A.
[0232] Additionally or alternatively, in response to the detection of the above relative to Figure 9B The described user input, computer system 101 displays one or more user interfaces for providing information about the 3D representation 916A and facilitating interaction with the 3D representation, such as Figure 9CAs shown. For example, in some embodiments, one or more user interfaces described above include a control user interface 920A for controlling the 3D representation 916A (e.g., corresponding to 916B in a top view of the 3D environment 902). In some embodiments, the control user interface 920A includes one or more optional options for facilitating user interaction with the 3D representation 916A. For example, the control user interface 920A includes a "close" button to terminate the display of the 3D representation 916A. Additionally or alternatively, the control user interface 920A includes a "play" button for animates the 3D representation 916A. In some embodiments, and in response to detecting user input pointing to the play button (e.g., an air pinch gesture when the user's attention is on the play button), the computer system 101 continuously rotates the 3D representation 916A through different perspective views. In some embodiments, the control user interface 920A includes an orientation information portion 918A for displaying a digital or other representation of the 3D orientation of the 3D representation 916A. In some embodiments, the orientation information portion 918A represents, digitally or otherwise (in degrees), the orientation of the three-dimensional representation 916A along the X, Y, and / or Z axes (e.g., three-dimensional). In some embodiments, the computer system facilitates user interaction with the orientation information portion 918A to change the orientation of the three-dimensional representation 916A. For example, the digital representation can be edited by the user, and in response to detecting that the digital representation has been edited by the user, the computer system 101 changes the orientation of the three-dimensional representation to correspond with the edited digital representation. In some embodiments, and as... Figure 9C As illustrated in the top view of the three-dimensional environment, a control user interface 920B, including an orientation information section 918B, is displayed in front of the three-dimensional representation 916A (e.g., closer to the user's viewpoint) and at a different location from the three-dimensional representation (e.g., relative to the user's viewpoint), allowing the user to view both the control user interface and the three-dimensional representation simultaneously. In some embodiments, when the three-dimensional representation 916A is displayed, the virtual object 908 (e.g., a two-dimensional representation on a virtual presentation) is not displayed. Additionally or alternatively, the two-dimensional representation may optionally be displayed and updated based on the displayed three-dimensional representation.
[0233] In some implementations, computer system 101 facilitates changes in the orientation of the displayed 3D representation 916A via user interaction with the 3D representation. For example, in response to user input (e.g., a pinch-in-the-air gesture when the user's attention (e.g., gaze 921) is directed at the 3D representation 916A and movement of the user's hand 903C while busy pinching in the air), the computer system changes the orientation of the 3D representation 916A, such as... Figure 9D exemplified. like Figure 9DAs illustrated, the orientation of the three-dimensional representation 916A has been modified by the computer system 101 to correspond to the movement of the user's hand 903C. For example, as Figure 9D As illustrated, the 3D representation 916A is now displayed from a bottom view. In some embodiments, in response to detecting a change in the orientation of a 3D object, the computer system 101 updates the orientation portion 918A to display a numerical representation of the new orientation of the 3D representation 916A that corresponds to the changed orientation. In some embodiments, the control user interface may include a "Restore" button that, when selected by the user, restores the orientation of the 3D representation 916A to its orientation prior to the user's change.
[0234] In some embodiments, the control user interface 920A includes an optional option 922 for terminating the display of the 3D representation 916A. In some embodiments, in response to detecting a selection of the optional option 922 (e.g., an air pinch gesture when the user's attention is directed to the optional option 922), the computer system terminates the display of the 3D representation 916A, such as... Figure 9E As illustrated. In some embodiments, and in response to a user's selection of optional option 922, the computing system displays the virtual object 908 (e.g., a two-dimensional representation of the virtual object) in the same orientation as the three-dimensional representation 916A was in when the user selected optional option 922. For example, and as shown... Figure 9E As illustrated, since the orientation of the 3D representation 916A terminates when it is displayed in the bottom perspective view, the virtual object 908 is now displayed from the bottom viewpoint. The virtual object 908 displayed on the virtual presentation 904A is shown at the same distance from the viewpoint as the virtual presentation 904A, compared to the 3D representation 916A.
[0235] Figure 10 This is a flowchart illustrating a method 1000 for displaying and controlling a three-dimensional representation of a virtual object in a demonstration application, according to some embodiments. In some embodiments, method 1000 is executed at a computer system (e.g., computer system 101 in Figure 1, such as a tablet computer, smartphone, wearable computer, or head-mounted device), which includes display generation components (e.g., Figure 1, ...). Figure 3 and Figure 4 The display generating component 120 (e.g., a heads-up display, monitor, touchscreen, and / or projector) and one or more cameras (e.g., cameras pointing downwards at the user's hand (e.g., color sensors, infrared sensors, or other depth-sensing cameras) or cameras pointing forward from the user's head). In some embodiments, method 1000 is performed by storing in a non-transitory computer-readable storage medium and by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., ...). Figure 1AThe control unit 110 in the middle executes instructions to manage. Some operations in method 1000 are optionally combined, and / or the order of some operations is optionally changed.
[0236] In some embodiments, method 1000 is performed at a computer system communicating with a display generation component and one or more input devices. In some embodiments, the computer system has one or more of the characteristics of the computer system in method 800. In some embodiments, the display generation component has one or more of the characteristics of the display generation component in method 800. In some embodiments, the one or more input devices have one or more of the characteristics of the one or more input devices in method 800.
[0237] In some implementations, when a virtual presentation associated with a demonstration application (e.g., as described in Reference Method 800) is displayed at a first location in a three-dimensional environment via a display generation component, the computer system receives (1002a) a first input pointing to a two-dimensional representation of a first three-dimensional virtual object via one or more input devices, wherein the virtual presentation includes the first three-dimensional virtual object (such as... Figure 9AThe virtual object (908) is a two-dimensional representation of the virtual object. In some embodiments, the two-dimensional representation of the first three-dimensional virtual object is displayed on a (virtual or physical) surface in a first three-dimensional environment associated with the displayed virtual presentation. In some embodiments, the two-dimensional representation is included and / or contained within virtual slides of the virtual presentation. Optionally, the virtual presentation is displayed on a two-dimensional surface in a three-dimensional environment or as a two-dimensional surface. In some embodiments, the virtual presentation can be edited by a user using a presentation application. For example, the computer system optionally moves the positioning of the two-dimensional representation of the virtual object on the virtual presentation when it receives an instruction to edit the virtual presentation. In some embodiments, the two-dimensional representation of the virtual object is based on the three-dimensional virtual object it is intended to represent. For example, the two-dimensional representation of the virtual object is an image of a three-dimensional virtual object rendered in two dimensions and corresponds to a view of the three-dimensional virtual object viewed from a particular perspective. In some embodiments, the three-dimensional environment has one or more of the characteristics of the first three-dimensional environment and / or the corresponding rehearsal virtual environment described in reference method 800. In some embodiments, an optional button is displayed along with the two-dimensional representation, and the first input includes the system detecting a selection of the optional button (e.g., via a tap or pinch on the optional button). In some embodiments, receiving the first input pointing to the two-dimensional representation includes detecting the user's gaze pointing to the two-dimensional representation. Additionally or alternatively, receiving the first input includes detecting an air pinch gesture pointing to a location in a three-dimensional environment associated with the two-dimensional representation. In some embodiments, while maintaining the air pinch, the computer system detects movement of the user's hand away from the virtual demo / object (e.g., to release the object). In some embodiments, the first input is received when the virtual demo is displayed in a demo mode of a demo application. The demo mode optionally displays the virtual demo in a non-editable form. Additionally or alternatively, the first input is received when the demo is displayed in a second mode different from the demo mode. For example, in some embodiments, the second mode includes displaying the virtual demo in a rehearsal virtual environment (e.g., as described with respect to method 800). Optionally, when the virtual demo is displayed in the second mode, the virtual demo can be edited by a user of the computing system via a demo application.
[0238] In some implementations, in response to receiving a first input, the computer system displays (1002b) a three-dimensional representation of a three-dimensional virtual object at a second location in a three-dimensional environment, different from the first location in the three-dimensional environment, wherein the second location is outside the virtual presentation, such as in Figure 9CIn some embodiments, displaying a 3D representation of a 3D virtual object includes simultaneously displaying a 2D representation and a 3D representation in a 3D environment. Optionally, displaying a 3D representation of a 3D virtual object includes stopping the computer system from displaying the 2D representation while the 3D representation is being displayed. In some embodiments, a second position for displaying the 3D representation in the 3D environment is outside and separate from the surface on which the virtual presentation is displayed in the 3D environment. Optionally, the second position for displaying the 3D representation of the virtual object in the 3D environment is in front of the surface on which the 2D representation of the virtual object is displayed in the virtual presentation. For example, from the user's perspective in 3D, the 3D representation of the virtual object appears to be in front of the virtual presentation and therefore closer to the user in the 3D environment than the virtual presentation (which includes the 2D representation of the virtual object) (e.g., closer to the user's viewpoint). Optionally, from the user's viewpoint, at least a portion of the virtual presentation is occluded by the 3D representation of the virtual object because the 3D object in front of the virtual presentation may block the user from seeing the occluded portion of the virtual presentation. Displaying the 2D representation of the 3D virtual object and rendering the 3D representation of the virtual object only when prompted by the user allows the device to minimize the amount of time it takes to display the 3D representation of the virtual object, thereby saving the computational resources of the computer system required to display the 3D object in a 3D environment.
[0239] In some implementations, a two-dimensional representation of the first three-dimensional virtual object in the virtual presentation is used in the virtual presentation along with a first visual indicator (such as...) to indicate that the virtual object is a three-dimensional virtual object. Figure 9B The visual indicator (910) is displayed together with the virtual representation in a three-dimensional environment. In some embodiments, the device displays a virtual representation of a three-dimensional environment on a two-dimensional surface. Therefore, the content of the visual representation is represented in two dimensions in the three-dimensional environment. Therefore, in order to distinguish the two-dimensional representation of a three-dimensional object on the virtual representation from ordinary two-dimensional objects displayed on the virtual representation, the device optionally displays a visual indicator along with the two-dimensional representation to indicate that the virtual object is viewable in three dimensions. In some embodiments, one or more visual indicators include text indicating that the virtual object is a three-dimensional object. Additionally or alternatively, one or more visual indicators include a box or other shape overlaid on the two-dimensional representation to indicate that the object is viewable in three dimensions. In some embodiments, the visual indicator is not displayed on the virtual representation unless the user's gaze is detected as pointing to the two-dimensional representation of the three-dimensional object on the virtual representation. Displaying a two-dimensional representation of a three-dimensional virtual object, displaying a visual indicator to indicate that the object is viewable in three dimensions, and rendering a three-dimensional representation of the virtual object only when prompted by the user allows the device to minimize the amount of time it takes to display the three-dimensional representation of the virtual object, thereby saving the computing resources of the computer system required to display three-dimensional objects in a three-dimensional environment.
[0240] In some implementations, when displaying a virtual presentation including a two-dimensional representation of a first three-dimensional virtual object, the computer system receives second input from a first portion of the user via one or more input devices. This second input includes a first air gesture pointing to a first visual indicator of the two-dimensional representation of the first three-dimensional virtual object, followed by movement of the user's first portion. The two-dimensional representation of the first three-dimensional virtual object is a first perspective view of the first three-dimensional virtual object, such as in… Figure 9B In some implementations, the second input includes detecting the user's attention (e.g., gaze) towards a two-dimensional representation of the virtual object, and detecting an air pinch when the user's attention is on the object, and then detecting movement of the user's hand in the shape of an air pinch. In some implementations, the direction / amount of the hand movement determines how the viewpoint of the virtual object is updated.
[0241] In some implementations, in response to receiving a second input, the computer system updates the two-dimensional representation within the virtual presentation based on detected movement of the user's first portion, wherein the updated two-dimensional representation of the first three-dimensional virtual object is a second perspective of the first three-dimensional virtual object, different from the first perspective, such as when the user rotates... Figure 9B In the case of virtual object 908 in virtual presentation 904, in some embodiments, the visual indicator used to identify the two-dimensional representation of the three-dimensional virtual object is interactive. When the device detects that a user has interacted with the visual indicator (e.g., by detecting a pinch and movement of the user's hand in the air), the device optionally rotates the two-dimensional representation of the three-dimensional virtual object according to the detected movement of the user's hand. In some embodiments, the two-dimensional representation is rotated using three-dimensional manipulation and represents a two-dimensional view of how the three-dimensional object looks from different perspectives. As will be described in further detail below, when the device displays the three-dimensional virtual object in a three-dimensional format, the device displays the object according to the current orientation of the object's two-dimensional representation on the presentation. Therefore, allowing the user to change the orientation of the two-dimensional representation of the virtual object ensures that the three-dimensional object is initially displayed in the user's preferred orientation. Allowing the user to rotate the two-dimensional representation of the virtual object minimizes the likelihood that the user will rotate the three-dimensional representation of the virtual object while it is being displayed, thereby saving computational and memory resources associated with rotating the three-dimensional object in a three-dimensional environment.
[0242] In some implementations, a two-dimensional representation of a first three-dimensional virtual object in the virtual presentation is displayed in the virtual presentation along with a second visual indicator used to initiate the process of displaying the three-dimensional representation of the first three-dimensional virtual object, such as in virtual object 708. Figure 7B In addition to Figure 9BIn cases where a visual indicator is included in addition to the visual indicator 910, in some embodiments, when displaying a virtual presentation including a two-dimensional representation of a first three-dimensional virtual object, the computer system receives a first input pointing to a second visual indicator via one or more input devices, the first input corresponding to a request to display a three-dimensional representation of the first three-dimensional virtual object. In some embodiments, the first input pointing to the visual indicator includes one or more characteristics of the input described above.
[0243] In some implementations, in response to receiving a first input pointing to a second visual indicator, the computer system displays a three-dimensional representation of a three-dimensional virtual object at a second location outside the virtual presentation, such as in Figure 9C In some embodiments, in addition to the visual indicator that identifies the two-dimensional representation of the virtual object as described above, the two-dimensional representation is also displayed along with a second visual indicator that is also interactive. When the computer system detects interaction with the second visual indicator, the device optionally renders a three-dimensional representation of the virtual object (described in further detail below). In some embodiments, the second visual indicator is not displayed until the computer system detects that the user's attention is directed at the object. In some embodiments, the second visual indicator includes, but is not limited to, a "play" symbol, text indicating 3D playback, or any other visual indicator designed to signal to the user that a three-dimensional representation of the virtual object will be rendered in a three-dimensional environment if the user interacts with the visual indicator. Displaying an interactive visual indicator to initiate the display of a three-dimensional representation of the three-dimensional virtual object minimizes the amount of user input required to display the three-dimensional representation, thereby saving computational resources associated with additional user input.
[0244] In some implementations, the second position is positioned in front of the virtual presentation relative to the user's viewpoint on the computer system, such as in... Figure 9C In some embodiments, the device displays a 3D representation of a virtual object in front of the virtual presentation from the user's viewpoint, in a manner that makes it appear as if the 3D representation of the object emerges from the virtual presentation. In some embodiments, the 3D representation occludes a portion of the virtual presentation during display because, from the user's perspective, the 3D representation is in front of the virtual presentation. Optionally, from the user's viewpoint, the 3D representation is displayed in front of the virtual presentation, but also on the side of the virtual presentation, such that no part of the 3D representation occludes any part of the virtual presentation. In some embodiments, the 3D representation is closer to the user's viewpoint than the virtual presentation. Displaying a 3D representation of the virtual object in front of it increases the visual salience of the 3D representation in the 3D environment, thereby reducing the likelihood of erroneous interaction with the 3D representation and thus conserving computational resources associated with correcting erroneous user input.
[0245] In some implementations, when a virtual presentation including a two-dimensional representation of a first three-dimensional virtual object is displayed and the second visual indicator is not displayed, the computer system receives a third input pointing to the two-dimensional representation via one or more input devices. In some implementations, in response to receiving the third input, the computer system displays the second visual indicator in the virtual presentation along with the two-dimensional representation of the first three-dimensional virtual object, such as when the user points the input to... Figure 9B In the case of virtual object 908, in some embodiments, when the computer system detects that the user has manipulated (e.g., via air pinch or other air gestures) a two-dimensional representation of the virtual object, the second visual indicator described above (e.g., an interactive visual indicator that initiates the display of a three-dimensional representation of the virtual object when selected or interacted with) appears together with the two-dimensional representation of the virtual object. In some embodiments, the second visual indicator appears when the device detects that the user's attention is directed at the two-dimensional representation of the virtual object (optionally, no further input from the user is required). In some embodiments, the second visual indicator is displayed for a predetermined amount of time after the computer system detects that the user has manipulated the two-dimensional representation. Once the device detects that the predetermined amount of time has expired, the device stops displaying the second visual indicator. The device will display the second visual indicator again when it detects that the user has manipulated the second visual indicator a second time. Displaying the second visual indicator only when the user manipulates the two-dimensional representation of the virtual object (which, when manipulated, initiates the display of a three-dimensional representation of the virtual object) preserves the computational resources associated with permanently displaying the second visual indicator on the visual presentation.
[0246] In some implementations, the two-dimensional representation of the three-dimensional virtual object is a first perspective view of the first three-dimensional virtual object. In some implementations, in response to receiving a first input, the computer system displays a three-dimensional representation of the three-dimensional virtual object from a first perspective relative to the user's current viewpoint, such as in... Figure 9C In some implementations, and as described above, the device allows the user to manipulate a two-dimensional representation of a virtual object to change its orientation in the virtual presentation. In some implementations, when the device initiates the display of a three-dimensional representation of a virtual object, the device displays the three-dimensional representation in the same orientation as the two-dimensional representation of the virtual object was in when the process of displaying the three-dimensional representation was initiated. In some implementations, if the two-dimensional representation comes from a second perspective (different from the first perspective), the three-dimensional representation will come from the second perspective instead of the first perspective. Displaying a three-dimensional representation of a virtual object in the same orientation as its two-dimensional representation in the virtual presentation minimizes the likelihood that the user will rotate the three-dimensional representation of the virtual object when it is displayed, thereby saving computational and memory resources associated with rotating a three-dimensional object in a three-dimensional environment.
[0247] In some implementations, when displaying a three-dimensional representation of a first three-dimensional virtual object, the computer system displays a control user interface (such as...) in the three-dimensional environment for controlling the three-dimensional representation at a third location within the three-dimensional environment. Figure 9C The control user interface (920A) is located in a third position, distinct from the first and second positions. In some embodiments, the device displays a control user interface (described in detail below) for controlling one or more aspects of the 3D representation, along with a 3D representation of the virtual object. In some embodiments, the control user interface is displayed in a position different from where the 3D representation is displayed, allowing the user to view both the control user interface and the 3D representation of the virtual object simultaneously without obscuring the view of the other. In some embodiments, the control user interface, the 3D representation, and the virtual presentation are at different or the same distance from the user's viewpoint. Displaying the control user interface in a position different from the 3D representation of the virtual object minimizes erroneous user input by allowing the user to view both simultaneously, thereby minimizing the likelihood of erroneous user input and saving computational resources associated with correcting erroneous user input.
[0248] In some implementations, when the control user interface is displayed, and after a predetermined time threshold has elapsed since the control user interface was displayed without any input directed to it, the computer system stops displaying the control user interface, such as in... Figure 9C In cases where the control user interface 920a stops displaying due to lack of input, in some embodiments, the device stops displaying the control user interface when it detects that the user has not interacted with the control user interface for longer than a predetermined time threshold. Optionally, the predetermined time threshold may be from 1 second to 1 hour. In some embodiments, when the predetermined time threshold has been exceeded, the device does not abruptly stop displaying the control user interface, but rather dims its display (e.g., gradually reduces the visual salience of the control user interface in a three-dimensional environment) until the control user interface is no longer visible in the three-dimensional environment. Displaying the control user interface for a predetermined time without user interaction conserves computing and memory resources that would otherwise be consumed in the case of persistent display of the control user interface.
[0249] In some implementations, when a 3D representation of a 3D virtual object is displayed, and when a control user interface is not currently displayed, the computer system receives a fourth input from the user via one or more input devices, while the user's attention is directed towards the 3D representation of the first 3D virtual object. In some implementations, in response to receiving the fourth input, a control user interface is displayed at a third location, such as... Figure 9CThe control user interface 920a is described above. In some embodiments, once the control user interface is no longer displayed by the device or has faded due to a lack of user interaction with the interface as described above, the device redisplays the control user interface in response to detecting input from the user. For example, the device detects an air pinch and / or user gaze at the location where the control user interface was previously displayed, and in response, redisplays the control user interface at the location where it was previously displayed in a three-dimensional environment. When the device detects user input pointing to the location where the control user interface was previously displayed, displaying the control user interface preserves computational and memory resources that would otherwise be consumed in the case of persistently displaying the control user interface.
[0250] In some embodiments, when a 3D representation of a first 3D virtual object is displayed, and when a control user interface is displayed, the computer system receives a fifth input directed to the control user interface via one or more input devices to stop displaying the 3D representation. In some embodiments, in response to receiving the fifth input, the 3D representation is stopped in a manner such as... Figure 9D The display in a three-dimensional environment. In some embodiments, the control user interface includes one or more optional options for turning off (e.g., stopping the display) the three-dimensional representation of the virtual object. In response to detecting a selection of one or more optional options, the device optionally stops the display of the three-dimensional representation and optionally redisplays the two-dimensional representation on the virtual presentation. In some embodiments, the device redisplays the two-dimensional representation of the virtual object in the same orientation as when the three-dimensional representation was turned off by the device, in response to user input at the control user interface. For example, if the viewpoint of the three-dimensional representation is a first viewpoint, then when the three-dimensional representation is turned off by the user, the two-dimensional representation of the virtual object on the virtual presentation is displayed in the first viewpoint. Similarly, if the viewpoint of the three-dimensional representation is a second viewpoint (different from the first viewpoint), then when the three-dimensional representation is turned off by the user, the two-dimensional representation of the virtual object on the virtual presentation is displayed in the second viewpoint. In some embodiments, the viewpoint of the two-dimensional representation is different from the original display viewpoint, for example, due to the user modifying the viewpoint. Stopping the display of the three-dimensional representation of the virtual object when receiving an instruction that the user no longer expects the three-dimensional representation to be displayed preserves computing and memory resources that would otherwise be consumed in the case of persistently displaying the control user interface.
[0251] In some implementations, when a 3D representation of a 3D object is displayed from a first-person perspective, the computer system receives instructions via one or more input devices to change the displayed viewpoint of the 3D representation, such as... Figure 9C Input 921 in the middle. In some implementations, in response to receiving an instruction to change the displayed viewpoint of the 3D representation, the computer system displays the 3D representation from a second viewpoint different from the first viewpoint, such as in Figure 9DIn some implementations, when a three-dimensional representation of a first three-dimensional object is displayed from a second perspective, the computer system receives a sixth input via one or more input devices, directing the user interface to change the perspective of the three-dimensional representation back to the default perspective. This input could include user selection of optional options on the user interface 920 to restore the orientation of the three-dimensional representation 916A to the default perspective. Figure 9D The previous state in the text.
[0252] In some implementations, in response to receiving a sixth input, the computer system displays a 3D representation of a 3D object from a default perspective, such as when the user control interface 920A includes an optional option to return the 3D representation 916a to its default orientation. In some implementations, the control user interface includes one or more optional options for restoring the orientation of the 3D object to the orientation in which the 3D representation of the virtual object was initially displayed when the display of the 3D representation was initiated. Therefore, the user can manipulate the orientation of the 3D object without needing to remember the original orientation of the 3D representation, allowing them to return the representation to its original orientation. Providing an optional option on the control user interface to display the 3D representation in its original orientation reduces the amount of user input required to modify the orientation of the 3D representation, thereby saving computational resources.
[0253] In some implementations, when displaying a 3D representation of a 3D virtual object, the computer system receives a seventh input via one or more input devices, directed to a control user interface, to animate the 3D representation, such as including... Figure 9D The control user interface 920a includes a "Play" button. In some embodiments, in response to receiving a seventh input, the computer system animates the 3D representation according to a predetermined animation sequence. In some embodiments, the control user interface includes one or more optional options for animates the 3D representation of the virtual object. In response to detecting a selection of one or more optional options, the device animates the 3D representation according to a predetermined animation sequence associated with the virtual object. In some embodiments, the animated 3D representation is capable of moving in a 3D manner (e.g., animated) within a 3D environment. Additionally or alternatively, the animated 3D representation remains at a fixed position in the 3D environment, a position commensurate with the 3D position when it is not animated. In some embodiments, the optional options for animates the 3D representation also serve to stop the animation. In some embodiments, the animated 3D representation progresses through the animation of different orientations of the virtual object. Providing optional options for animates and stopping the animation on the control user interface saves computational and memory resources associated with persistently animates the 3D representation in a 3D environment.
[0254] In some implementations, the control user interface includes one or more orientation information portions for indicating the orientation of a 3D representation of a 3D object relative to the 3D environment, such as... Figure 9D The orientation information section 918a is included. In some embodiments, the device displays one or more numerical representations of the orientation of a three-dimensional representation of a virtual object as part of a control user interface. For example, one or more numerical representations include rotation degrees on each of the X, Y, and Z axes. In some embodiments, the device updates one or more numerical representations in response to a change in the orientation of the three-dimensional representation. Providing a numerical representation of the orientation of a three-dimensional object relative to the three-dimensional environment minimizes erroneous user input associated with rotating the three-dimensional object to a particular orientation, thereby saving computational resources associated with correcting erroneous user input.
[0255] In some implementations, when a control user interface including one or more orientation information portions is displayed, the computer system receives an eighth input pointing to one or more orientation information portions via one or more input devices. In some implementations, in response to receiving the eighth input, the computer system modifies the orientation of the 3D representation relative to the 3D environment based on the received eighth input, such as when the user points the input to... Figure 9D The orientation portion 918a in the middle is modified. Figure 9D In the case of the orientation of the three-dimensional representation 916a in the above embodiment, in some implementations, one or more numerical representations (described above) can be selected and modified by the user. In response to detecting that the user has selected and modified one or more numerical representations, the device changes the orientation of the three-dimensional representation of the virtual object according to the modified numerical representation. Allowing the user to modify the numerical representation allows the user to modify the orientation of the three-dimensional representation of the virtual object more precisely, thereby minimizing erroneous modifications to the orientation and thus preserving the computational resources associated with correcting erroneous modifications to the orientation.
[0256] In some embodiments, when displaying a 3D representation of a first 3D virtual object, the computer system receives a ninth input from a second part of the user via one or more input devices, the ninth input including a second air gesture pointing at the 3D representation, followed by movement of the user's second part. In some embodiments, in response to receiving the ninth input, the computer system rotates the 3D representation relative to the 3D environment based on detected movement of the user's second part, such as when the user points the input... Figure 9CIn cases where the orientation of a 3D representation 916A is changed by manipulating the representation, in some embodiments, the device modifies (e.g., rotates) the orientation of the 3D representation of a virtual object in response to detecting the user's attention and an air gesture (e.g., an air pinch) pointing at the 3D representation, and subsequently detecting movement of a part of the user's body (e.g., the user's hand, while maintaining the air pinch). In some embodiments, the amount of rotation corresponds to the detected amount of hand movement, the axis of rotation corresponds to the direction of hand movement, and the rotation speed corresponds to the detected speed of hand movement. In some embodiments, the 3D representation is rotated according to the detected direction of movement of that part of the user's body. Allowing the user to modify the orientation of the 3D representation of a virtual object by manipulating the representation can result in a more accurate modification of the orientation of the representation, thereby minimizing erroneous modifications to the orientation and thus conserving computational resources associated with correcting erroneous modifications to the orientation.
[0257] In some implementations, the 3D representation of the 3D virtual object is displayed when the demo application is in edit mode, for example, when the demo application is in... Figures 9A to 9E In the context of the edit mode, in some implementations, "edit mode" refers to a mode of the presentation application in which the user can edit the content of the virtual presentation. In some implementations, and when in edit mode, the device displays a three-dimensional representation of the virtual object according to the implementations described above. For example, the three-dimensional representation is displayed by the device in response to one or more user inputs as described above. In some implementations, and when displaying the three-dimensional representation, the device also displays a control user interface when the presentation application is in edit mode. In some implementations, in edit mode, the user can use one or more tools to edit the presentation. For example, in edit mode, the user can add and / or delete content and modify existing content. Displaying a three-dimensional representation when the presentation application is in edit mode allows the user to accurately edit the virtual presentation, thereby minimizing erroneous user input associated with inaccurate modifications to the virtual presentation and thus saving computational resources associated with correcting erroneous input.
[0258] In some implementations, the 3D representation of the 3D virtual object is displayed when the demo application is in demo mode, for example, when the demo application is in Figures 9A to 9EIn the context of the demonstration mode, in some implementations, "demonstration mode" refers to a mode in which the device displays a virtual demo and the user cannot edit the content of the virtual demo (e.g., the user cannot add and / or delete content, and / or modify existing content). In some implementations, and when in demonstration mode, the device displays a three-dimensional representation of the virtual object according to the implementations described above. For example, the three-dimensional representation is displayed by the device in response to one or more user inputs as described above. In some implementations, and when displaying the three-dimensional representation, the device also displays a control user interface when the demonstration application is in edit mode. Displaying a three-dimensional representation when the demonstration application is in demonstration mode allows the user to accurately view the virtual demo, thereby minimizing erroneous user input associated with inaccurate modifications to the virtual demo, and thus saving computational resources associated with correcting erroneous input.
[0259] In some implementations, a first viewpoint position is determined based on the user's viewpoint of the computer system at the time the first input is detected, and a second position in the three-dimensional environment is a first corresponding position. In some implementations, a second viewpoint position is determined based on the user's viewpoint of the computer system at the time the first input is detected, which is different from the first viewpoint position, and a second position in the three-dimensional environment is a second corresponding position different from the first corresponding position, such as when the virtual presentation 904A is displayed in the presenter's or audience's viewpoint as described above with respect to method 800. In some implementations, the position where the three-dimensional representation is displayed depends on whether the virtual presentation is displayed from the audience's viewpoint or the presenter's viewpoint. In some implementations, if the device displays the virtual presentation from the audience's viewpoint (described above with respect to method 800), the three-dimensional representation of the virtual object is displayed in front of the virtual presentation and in front of the user from the user's perspective. In some implementations, if the device displays the virtual presentation from the presenter's viewpoint (described above with respect to method 800), the three-dimensional representation of the virtual object is displayed in front of the virtual presentation, but behind the user from the user's perspective. In some implementations, the first corresponding position is based on the positioning of the user's viewpoint in the three-dimensional environment. For example, if the user is off-center and on the left side of the virtual presentation, the first corresponding position will also be off-center and on the left side of the virtual presentation; similarly, if the user is off-center and on the right side of the virtual presentation, the corresponding position will also be off-center and on the right side of the virtual presentation. Displaying the 3D representation according to the virtual presentation's viewpoint minimizes the amount of user input required to place the 3D representation within the 3D environment, thus saving computational resources.
[0260] Figures 11A to 11D Examples of audio models demonstrated within one or more virtual environments associated with a demonstration application, according to some implementation schemes, are illustrated.
[0261] Figure 11AA top-down view of a 3D environment 1102 associated with a demonstration application is shown, in which audio is being presented based on an audio model (described in detail below). For illustrative purposes, Figure 11A The illustrated three-dimensional environment 1102 represents the above relative to Figure 7E The auditorium rehearsal virtual environment described (and in contrast to that described in method 800), but in contrast to Figures 11A to 11D The described concepts can be applied to any three-dimensional environment associated with a demonstration application. In some implementations, the computer system displays the three-dimensional environment from a "presenter's viewpoint" (as described in relation to method 800). Therefore, the virtual presentation 1112 (described above with respect to methods 800 and 1000) is positioned behind the user 1104, while the audience is positioned in front of the user, such as... Figure 11A exemplified.
[0262] In some embodiments, and to simulate a real-world auditorium environment, the three-dimensional environment 1102 includes one or more virtual audio speakers 1106A to 1106D. Based on the physical locations where real-world audio speakers would be placed in a real-world auditorium, the virtual audio speakers 1106A to 1106D are optionally placed in the three-dimensional environment. In some embodiments, the device displays audio in the three-dimensional environment as if the audio were emitted from a location in the three-dimensional environment 1102 associated with the virtual audio speakers 1106A to 1106D, without displaying any representation of the virtual audio speakers. In some embodiments, the device displays representations of the virtual audio speakers 1106A to 1106D within the three-dimensional environment 1102 (in addition to displaying audio as if the audio were emitted from a location associated with the virtual audio speakers 1106A to 1106D).
[0263] In some embodiments, virtual speakers 1106A to 1106D emit audio associated with the virtual presentation 1112. For example, if the virtual presentation 1112 includes audio content, the audio content is optionally emitted from the virtual speakers 1106A to 1106D. Additionally or alternatively, the computer system may optionally emit audio associated with a user (e.g., the presenter of the virtual presentation) via the virtual speakers 1106A to 1106D. For example, in a real-world auditorium environment, a presenter speaks to an audience through a microphone, which is then amplified and used to present the audio to the audience via speakers placed throughout the auditorium. Therefore, in some embodiments, and to simulate a real-world auditorium environment, the computer system optionally collects audio from the user (via a microphone or other audio collection device) and transmits the collected audio through the virtual audio speakers 1106A to 1106D.
[0264] In one or more examples, the computer system demonstrates audio emitted by user 1104 and audio emitted at virtual audio speakers 1106A to 1106D, based on audio model 1110. In some embodiments, the audio model refers to one or more characteristics applied to the audio demonstrated by the computer system in a three-dimensional environment 1102. In some embodiments, the audio characteristics associated with the audio model include spatial audio characteristics. Examples of spatial audio characteristics optionally include, but are not limited to, audio characteristics associated with the directionality of sound, reverberation (e.g., echo), pitch, etc. For example, as Figure 11A As illustrated, audio from user 1104 and / or virtual audio speakers 1106a to 1106d bounces off objects represented in a three-dimensional environment. For example, relative to... Figure 11A The illustrated auditorium virtual environment features audio reflected from the walls of the environment, as illustrated at 1108. Additionally or alternatively, audio may be reflected from objects in the environment, such as chairs or other objects found within the virtual environment. The spatial audio characteristics of the audio model may optionally incorporate the spatial features of the three-dimensional environment 1102 described above. In some embodiments, the audio characteristics associated with the audio model include environmental audio characteristics. Examples of environmental audio characteristics include white noise or other noise associated with the real-world environment that the three-dimensional environment 1102 is intended to mimic. In some embodiments, the environmental audio characteristics are based on recordings of environmental noise from the physical environment corresponding to the three-dimensional environment 1102.
[0265] In some implementations, the audio model 1110 includes one or more audio parameters that are associated with both spatial audio characteristics and environmental audio characteristics (and with other audio characteristics of the audio model). Figure 11A As illustrated, a sliding rule is used to represent one or more audio parameters of audio model 1110 to illustrate the value associated with each parameter, so as to represent the value of each parameter that the computer system will set when demonstrating audio according to the audio model. In some embodiments, demonstrating audio according to the audio model encompasses the computer system setting the value of each audio parameter to a predetermined value that is commensurate with the audio characteristics (e.g., space and environment) associated with a particular three-dimensional environment. Thus, audio model 1110 is optionally based on the three-dimensional environment 1102 in which the virtual demonstration is being performed. As described below, in response to changes in the three-dimensional environment, the computer system optionally modifies the audio model used to demonstrate audio in the three-dimensional environment.
[0266] Figure 11B Examples are given of applications when the device displays the environment from different viewpoints. Figure 11A The audio model of the three-dimensional environment. As described above relative to method 800, the device can display different viewpoints of the three-dimensional environment associated with the demonstration application. For example, instead of in Figure 11AThe device displays a 3D environment 1102 (described above relative to method 800) from the "presenter's" viewpoint, and from the "audience's" viewpoint. In some embodiments, the dev...
Claims
1. A method, the method comprising: At the computer system that communicates with the display generation component and one or more input devices: When a virtual presentation associated with a presentation application is displayed in a first three-dimensional environment via the display generation component, a first input corresponding to a request to display the virtual presentation in a second mode different from a first mode of the presentation application is received via the one or more input devices, wherein the first three-dimensional environment includes a portion of the physical environment of the user of the computer system, and the virtual presentation is presented in the first mode of the presentation application; as well as In response to receiving the first input: The process of initiating the display of the virtual presentation in a corresponding virtual environment different from the first three-dimensional environment is described, wherein when the virtual presentation is displayed in the corresponding virtual environment, the portion of the user's physical environment is invisible via the display generation component.
2. The method according to claim 1, further comprising: In response to receiving the first input: The display generation component displays a virtual environment selection user interface for selecting one or more predetermined virtual environments of the demo application; when the virtual environment selection user interface is displayed, a second input corresponding to the selection of a virtual environment in the one or more predetermined virtual environments is received via the one or more input devices; as well as In response to receiving the second input, the virtual presentation is displayed in a selected virtual environment among the one or more virtual environments.
3. The method of claim 2, wherein displaying the virtual presentation in the selected, predetermined virtual environment comprises: The selected predefined virtual environment is displayed from a viewpoint corresponding to the presenter's viewpoint within the selected predefined virtual environment.
4. The method according to any one of claims 2 to 3, wherein displaying the virtual presentation in the selected, predetermined virtual environment in response to the second input comprises: The display generation component displays a virtual environment settings user interface for changing one or more settings associated with the selected, predetermined virtual environment.
5. The method according to claim 4, further comprising: When the virtual environment settings user interface is displayed, a third input directed to the virtual environment settings user interface is received via the one or more input devices to change one or more virtual light settings associated with the selected predetermined virtual environment; And in response to receiving the third input for changing the one or more virtual light settings, adjust the virtual lighting characteristics of the selected predetermined virtual environment according to the third input.
6. The method according to any one of claims 4 to 5, further comprising: When the virtual environment settings user interface is displayed, a third input corresponding to a request to change the viewpoint of the selected, predetermined virtual environment to the presenter's viewpoint is received via the one or more input devices; as well as In response to receiving the third input, and based on determining that the virtual environment is displayed from a viewpoint different from the presenter's viewpoint, displaying the selected predetermined virtual environment from the presenter's viewpoint, wherein displaying the selected predetermined virtual environment from the presenter's viewpoint includes: placing the virtual presentation behind the user's viewpoint within the selected virtual environment.
7. The method according to any one of claims 4 to 6, further comprising: When the virtual environment settings user interface is displayed, a fourth input corresponding to a request to change the viewpoint of the selected, predetermined virtual environment to the viewpoint of the viewer is received via the one or more input devices; and In response to receiving the fourth input, and based on determining that the virtual environment is displayed from a viewpoint different from the viewpoint of the user, displaying the selected predetermined virtual environment from the viewpoint of the user, wherein displaying the selected predetermined virtual environment from the viewpoint of the user includes: displaying the virtual presentation in the selected virtual environment in front of the user's viewpoint.
8. The method according to any one of claims 4 to 7, wherein the virtual environment settings user interface includes one or more interactive options, the one or more interactive options being interactive in response to detecting input from a first part of the user when the user's attention is directed to an optional option.
9. The method according to any one of claims 4 to 8, further comprising: When the virtual settings user interface is displayed in an expanded state, a third input directed to the virtual settings user interface is received via the one or more input devices, wherein the third input includes an input from a first part of the user's body directed to a first part of the virtual settings user interface, followed by a downward movement of the first part of the user's body.
10. The method according to any one of claims 4 to 9, further comprising: When the virtual environment settings user interface is displayed, a third input is received via the one or more input devices to activate the virtual laser pointer in the selected virtual environment; In response to receiving the third input for activating the virtual laser pointer, a virtual laser point is displayed via the display generation component, wherein the virtual laser pointer moves within a selected virtual environment based on detected movement of a first part of the user.
11. The method according to claim 10, further comprising: Receive a fourth input via the one or more input devices, the fourth input including an air gesture from the user’s first part pointing at the virtual demonstration; as well as In response to the detection of the fourth input: Based on the determination that the virtual laser pointer is not activated, navigate in the virtual demonstration according to the fourth input; as well as Based on the determination that the virtual laser point is activated, navigation in the virtual demonstration is abandoned based on the fourth input.
12. The method according to any one of claims 1 to 11, wherein displaying the virtual presentation in the corresponding virtual environment comprises: Based on the determination that the virtual presentation is in presentation mode, a speaker notes user interface for displaying information associated with the virtual presentation is displayed via the display generation component.
13. The method of claim 12, wherein the speaker notes user interface is displayed in the corresponding virtual environment at a location different from the location where the virtual presentation is displayed in the corresponding virtual environment.
14. The method according to any one of claims 12 to 13, further comprising: When the speaker notes user interface is displayed, a second input is received from a first part of the user via the one or more input devices, the second input including a first air pinch pointing to a part of the speaker notes user interface, followed by movement of the first part of the user; as well as In response to receiving the second input, the speaker notes user interface is moved within the corresponding virtual environment based on the detected movement of the user's first part.
15. The method according to any one of claims 1 to 14, the method further comprising: When the virtual demo is displayed, a second input is received via the one or more input devices from a part of the user's body, the second input including a first air gesture from the user's first part of the body, the first air gesture including movement of the first part of the body in a corresponding direction; and In response to receiving the second input: Based on determining that the corresponding direction is a first direction, navigate in the virtual demonstration in the first corresponding direction; as well as Based on the determination that the corresponding direction is a second direction different from the first direction, navigation is performed in the virtual demonstration in the second corresponding direction different from the first corresponding direction.
16. The method of claim 15, wherein the first input points to a speaker notes user interface that is displayed concurrently with the virtual presentation in the corresponding virtual environment.
17. The method of claim 15, wherein the first input is directed to the virtual presentation in the corresponding virtual environment.
18. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a virtual presentation associated with a presentation application is displayed in a first three-dimensional environment via the display generation component, a first input corresponding to a request to display the virtual presentation in a second mode different from a first mode of the presentation application is received via the one or more input devices, wherein the first three-dimensional environment includes a portion of the physical environment of the user of the computer system, and the virtual presentation is presented in the first mode of the presentation application; as well as In response to receiving the first input: The process of initiating the display of the virtual presentation in a corresponding virtual environment different from the first three-dimensional environment is described, wherein when the virtual presentation is displayed in the corresponding virtual environment, the portion of the user's physical environment is invisible via the display generation component.
19. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; A means for receiving, via the one or more input devices, a first input corresponding to a request to display the virtual presentation in a second mode different from a first mode of the presentation application when a virtual presentation associated with a presentation application is displayed in a first three-dimensional environment via the display generation component, wherein the first three-dimensional environment includes a portion of the physical environment of the user of the computer system, and the virtual presentation is presented in the first mode of the presentation application. and A means for performing the following operations in response to receiving the first input: The process of initiating the display of the virtual presentation in a corresponding virtual environment different from the first three-dimensional environment is described, wherein when the virtual presentation is displayed in the corresponding virtual environment, the portion of the user's physical environment is invisible via the display generation component.
20. An electronic device, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 1 to 17.
21. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any one of the methods according to claims 1 to 17.
22. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; and Apparatus for performing any one of the methods according to claims 1 to 17.
23. A method, the method comprising: At the computer system that communicates with the display generation component and one or more input devices: When a virtual demo associated with a demo application is displayed at a first location in a three-dimensional environment via the display generation component, a first input pointing to a two-dimensional representation of a first three-dimensional virtual object is received via the one or more input devices, wherein the virtual demo includes the two-dimensional representation of the first three-dimensional virtual object; as well as In response to receiving the first input, a three-dimensional representation of the three-dimensional virtual object is displayed at a second location in the three-dimensional environment that is different from the first location in the three-dimensional environment, wherein the second location is outside the virtual presentation.
24. The method of claim 23, wherein the two-dimensional representation of the first three-dimensional virtual object in the virtual demonstration is displayed in the virtual demonstration together with a first visual indicator for indicating that the virtual object is a three-dimensional virtual object.
25. The method according to claim 24, further comprising: When the virtual presentation including the two-dimensional representation of the first three-dimensional virtual object is displayed, a second input from a first part of the user is received via the one or more input devices. The second input includes a first air gesture pointing to the first visual indicator of the two-dimensional representation of the first three-dimensional virtual object, followed by movement of the first part of the user. The two-dimensional representation of the first three-dimensional virtual object is a first perspective view of the first three-dimensional virtual object. as well as In response to receiving the second input, the two-dimensional representation within the virtual presentation is updated based on the detected movement of the first part of the user, wherein the updated two-dimensional representation of the first three-dimensional virtual object is a second perspective of the first three-dimensional virtual object that is different from the first perspective.
26. The method of any one of claims 23 to 25, wherein the two-dimensional representation of the first three-dimensional virtual object of the virtual demonstration is displayed in the virtual demonstration together with a second visual indicator for initiating the process of displaying the three-dimensional representation of the first three-dimensional virtual object, the method further comprising: When the virtual presentation including the two-dimensional representation of the first three-dimensional virtual object is displayed, the first input pointing to the second visual indicator is received via the one or more input devices, the first input corresponding to a request to display the three-dimensional representation of the first three-dimensional virtual object; as well as In response to receiving the first input pointing to the second visual indicator, the three-dimensional representation of the three-dimensional virtual object is displayed at a second location outside the virtual presentation.
27. The method of claim 26, wherein the second position is in front of the virtual presentation relative to the viewpoint of the user of the computer system.
28. The method according to any one of claims 26 to 27, further comprising: When the virtual demonstration including the two-dimensional representation of the first three-dimensional virtual object is displayed and the second visual indicator is not displayed, a third input pointing to the two-dimensional representation is received via the one or more input devices; as well as In response to receiving the third input, the second visual indicator is displayed in the virtual presentation together with the two-dimensional representation of the first three-dimensional virtual object.
29. The method according to any one of claims 26 to 28, wherein the two-dimensional representation of the three-dimensional virtual object is a first perspective view of the first three-dimensional virtual object, and wherein the method further comprises: In response to receiving the first input, the three-dimensional representation of the three-dimensional virtual object is displayed from a first perspective relative to the user's current viewpoint.
30. The method according to any one of claims 23 to 29, further comprising: When the three-dimensional representation of the first three-dimensional virtual object is displayed, a control user interface for controlling the three-dimensional representation at a third location in the three-dimensional environment is displayed, wherein the third location is different from the first location and the second location.
31. The method according to claim 30, further comprising: When the control user interface is displayed, and after the control user interface has been displayed for longer than a predetermined time threshold without any input to it, the display of the control user interface is stopped.
32. The method according to claim 31, further comprising: When the three-dimensional representation of the three-dimensional virtual object is displayed, and when the control user interface is not being displayed, a fourth input from the first part of the user is received via the one or more input devices when the user's attention is directed at the three-dimensional representation of the first three-dimensional virtual object; as well as In response to receiving the fourth input, the control user interface is displayed at the third location.
33. The method according to any one of claims 30 to 32, further comprising: When the three-dimensional representation of the first three-dimensional virtual object is displayed, and when the control user interface is displayed, a fifth input directed to the control user interface is received via the one or more input devices to stop the display of the three-dimensional representation; as well as In response to receiving the fifth input, the display of the three-dimensional representation in the three-dimensional environment is stopped.
34. The method according to any one of claims 30 to 33, further comprising: When the three-dimensional representation of the three-dimensional object is displayed from a first perspective, an instruction to change the display perspective of the three-dimensional representation is received via the one or more input devices; In response to receiving the instruction to change the display perspective of the three-dimensional representation, the three-dimensional representation is displayed from a second perspective different from the first perspective; When the three-dimensional representation of the first three-dimensional object is displayed from the second perspective, a sixth input directed to the control user interface is received via the one or more input devices to change the perspective of the three-dimensional representation back to the default perspective; as well as In response to receiving the sixth input, the three-dimensional representation of the three-dimensional object is displayed from the default perspective.
35. The method according to any one of claims 30 to 34, further comprising: When the three-dimensional representation of the three-dimensional virtual object is displayed, a seventh input directed to the control user interface is received via the one or more input devices to animate the three-dimensional representation; and In response to receiving the seventh input, the 3D representation is animated according to a predetermined animation sequence.
36. The method according to any one of claims 30 to 35, wherein the control user interface includes one or more orientation information portions for indicating the orientation of the three-dimensional representation of the three-dimensional object relative to the three-dimensional environment.
37. The method according to claim 36, further comprising: When a control user interface including the one or more orientation information portions is displayed, an eighth input pointing to the one or more orientation information portions is received via the one or more input devices; as well as In response to receiving the eighth input, the orientation of the three-dimensional representation relative to the three-dimensional environment is modified according to the received eighth input.
38. The method according to any one of claims 23 to 37, further comprising: When the three-dimensional representation of the first three-dimensional virtual object is displayed, a ninth input from the user's second part is received via the one or more input devices, the ninth input including a second air gesture pointing at the three-dimensional representation, followed by movement of the user's second part; as well as In response to receiving the ninth input, the three-dimensional representation is rotated relative to the three-dimensional environment based on the detected movement of the second part of the user.
39. The method according to any one of claims 23 to 38, wherein the three-dimensional representation of the three-dimensional virtual object is displayed when the demo application is in edit mode.
40. The method according to any one of claims 23 to 39, wherein the three-dimensional representation of the three-dimensional virtual object is displayed when the demo application is in demo mode.
41. The method according to any one of claims 23 to 40, wherein: Based on the determination that the user's viewpoint in the computer system when the first input is detected is the first viewpoint position, the second position in the three-dimensional environment is the first corresponding position; and Based on the determination that when the first input is detected, the user's viewpoint in the computer system is a second viewpoint position different from the first viewpoint position, and the second position in the three-dimensional environment is a second corresponding position different from the first corresponding position.
42. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a virtual demo associated with a demo application is displayed at a first location in a three-dimensional environment via the display generation component, a first input pointing to a two-dimensional representation of a first three-dimensional virtual object is received via the one or more input devices, wherein the virtual demo includes the two-dimensional representation of the first three-dimensional virtual object; as well as In response to receiving the first input, a three-dimensional representation of the three-dimensional virtual object is displayed at a second location in the three-dimensional environment that is different from the first location in the three-dimensional environment, wherein the second location is outside the virtual presentation.
43. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; A means for receiving, via the one or more input devices, a first input pointing to a two-dimensional representation of a first three-dimensional virtual object when a virtual presentation associated with a demonstration application is displayed at a first location in a three-dimensional environment via the display generation component, wherein the virtual presentation includes the two-dimensional representation of the first three-dimensional virtual object; and A means for displaying a three-dimensional representation of a three-dimensional virtual object at a second location in the three-dimensional environment, different from the first location in the three-dimensional environment, in response to receiving the first input, wherein the second location is outside the virtual representation.
44. An electronic device, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 23 to 41.
45. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any one of the methods according to claims 23 to 41.
46. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; and Apparatus for performing any one of the methods according to claims 23 to 41.
47. A method comprising: At the computer system that communicates with the display generation component and one or more input devices: When a virtual demo associated with a demo application is displayed in a first three-dimensional environment via the display generation component, a first input corresponding to a request to display the virtual demo in a second three-dimensional environment is received via the one or more input devices, the virtual demo including audio corresponding to the virtual demo based on a first audio model associated with the first three-dimensional environment; as well as In response to receiving the first input, the virtual presentation is displayed in the second three-dimensional environment, the virtual presentation including the presentation of audio corresponding to the virtual presentation based on a second audio model that is different from the first audio model associated with the second three-dimensional environment.
48. The method of claim 47, wherein the first audio model comprises one or more first spatial audio parameters, wherein the second audio model comprises one or more second spatial audio parameters different from the one or more first spatial audio parameters, and wherein the method further comprises: In response to receiving the first input, audio corresponding to the virtual demonstration associated with the second three-dimensional environment is displayed based on the one or more second spatial audio parameters.
49. The method of claim 48, wherein the one or more first spatial audio parameters include one or more first reverberation parameters, wherein the one or more second spatial audio parameters include one or more second reverberation parameters different from the one or more first reverberation parameters, and wherein the method further comprises: In response to receiving the first input, audio corresponding to the virtual demonstration associated with the second three-dimensional environment is presented based on the one or more second reverberation parameters.
50. The method of any one of claims 48 to 49, wherein the one or more second spatial audio parameters are based on a physical environment associated with the second three-dimensional environment.
51. The method according to any one of claims 47 to 50, wherein the first audio model comprises a first ambient audio model, wherein the second audio model comprises a second ambient audio model different from the first ambient audio model, and wherein the method further comprises: In response to receiving the first input, audio corresponding to the virtual demonstration associated with the second three-dimensional environment is displayed according to the second environmental audio model.
52. The method of claim 51, wherein the second environment model is based on one or more audio recordings of a physical environment associated with a second three-dimensional environment.
53. The method according to any one of claims 51 to 52, wherein demonstrating audio corresponding to the virtual demonstration associated with the second three-dimensional environment according to the second environmental audio model comprises: Based on the determination that the user's viewpoint in the second three-dimensional environment of the computer system is the first viewpoint, the first audio is demonstrated; as well as Based on the determination that the user's viewpoint in the second three-dimensional environment of the computer system is a second viewpoint different from the first viewpoint, a second audio different from the first audio is demonstrated.
54. The method of claim 53, wherein the viewpoint of the second three-dimensional environment is a presenter's viewpoint.
55. The method of claim 53, wherein the viewpoint of the second three-dimensional environment is the viewpoint of the viewer.
56. The method according to any one of claims 47 to 55, wherein the second three-dimensional environment includes a first virtual audio speaker at a first location in the second three-dimensional environment, and wherein the method further comprises: In response to receiving the first input, an audio signal associated with the user of the computer system is output as if it were emitted from the first virtual audio speaker located at the first position in the second three-dimensional environment.
57. The method of claim 56, wherein the first position of the first virtual audio speaker in the second three-dimensional environment is based on the position of the physical audio speaker in the physical environment associated with the second three-dimensional environment.
58. The method of any one of claims 56 to 57, wherein outputting audio associated with the user of the computer system as if emitted from the first virtual audio speaker comprises: Output an amplified version of the audio detected by the user.
59. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a virtual demo associated with a demo application is displayed in a first three-dimensional environment via the display generation component, a first input corresponding to a request to display the virtual demo in a second three-dimensional environment is received via the one or more input devices, the virtual demo including audio corresponding to the virtual demo based on a first audio model associated with the first three-dimensional environment; as well as In response to receiving the first input, the virtual presentation is displayed in the second three-dimensional environment, the virtual presentation including the presentation of audio corresponding to the virtual presentation based on a second audio model that is different from the first audio model associated with the second three-dimensional environment.
60. An electronic device in communication with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; A means for receiving, via the one or more input devices, a first input corresponding to a request to display the virtual presentation in a second three-dimensional environment when a virtual presentation associated with a presentation application is displayed in a first three-dimensional environment via the display generation component, the virtual presentation including audio corresponding to the virtual presentation based on a first audio model associated with the first three-dimensional environment; and A means for displaying the virtual presentation in a second three-dimensional environment in response to receiving the first input, the virtual presentation including displaying audio corresponding to the virtual presentation according to a second audio model different from the first audio model associated with the second three-dimensional environment.
61. An electronic device, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 47 to 58.
62. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any one of the methods according to claims 47 to 58.
63. An electronic device communicating with a display generating component and one or more input devices, the electronic device comprising: One or more processors; Memory; and Apparatus for performing any one of the methods according to claims 47 to 58.