Devices, methods, and graphical user interfaces for interacting with a three-dimensional environment
The system addresses inefficiencies in virtual and augmented reality interactions by adjusting immersion levels and prioritizing physical objects, resulting in a more efficient and intuitive user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-11-08
- Publication Date
- 2026-04-15
AI Technical Summary
Existing methods and interfaces for interacting with virtual and augmented reality environments are cumbersome, inefficient, and complex, leading to increased cognitive burden and energy consumption, particularly in battery-powered devices.
The system enhances interaction by displaying three-dimensional computer-generated environments with adjustable immersion levels, prioritizing important physical objects, and using input devices like cameras and touch-sensitive surfaces to simplify user interaction.
This approach reduces user input errors, enhances efficiency, and improves the user experience by aligning device responses with user inputs, thus creating a more intuitive and energy-efficient human-machine interface.
Smart Images

Figure 0007846736000001 
Figure 0007846736000002 
Figure 0007846736000003
Abstract
Description
Technical Field
[0001] (Related Applications) This application is a continuation of U.S. patent application Ser. No. 17 / 483,722, filed Sep. 23, 2021, which claims priority to U.S. Provisional Patent Application No. 63 / 082,933, filed Sep. 24, 2020, the entire contents of each of which are incorporated herein by reference.
[0002] (Technical Field) The present disclosure generally relates to a computer system having a display generation component, including but not limited to, an electronic device that provides virtual and mixed reality experiences via a display, and one or more input devices that provide a computer-generated reality (CGR) experience.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include digital images, videos, text, icons, and virtual objects such as buttons and other graphics.
[0004] However, the methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are often cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems where manipulating virtual objects is complex and error-prone impair the user's cognitive burden and detract from the virtual / augmented reality experience. In addition, these methods are unnecessarily time-consuming and thus waste energy. This latter consideration is particularly important in battery-powered devices. [Overview of the project]
[0005] Therefore, there is a need for computer systems with improved methods and interfaces to provide users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. The above-mentioned deficiencies and other problems related to user interfaces for computer systems having display generation components and one or more input devices are mitigated or eliminated by the disclosed systems, methods, and user interfaces. Such systems, methods, and interfaces optionally complement or replace conventional systems, methods, and user interfaces for providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or types of input from the user by helping the user understand the connection between the inputs provided and the device responses to those inputs, thereby generating a more efficient human-machine interface.
[0006] According to some embodiments, the method is performed by a computer system that communicates with a first display generation component, one or more audio output devices, and one or more input devices, and includes: displaying a three-dimensional computer-generated environment via the first display generation component; detecting a first event corresponding to a request to present first computer-generated content, which includes first visual content and first audio content corresponding to the first visual content, while the three-dimensional computer-generated environment is being displayed; and in response to the detection of the first event corresponding to a request to present first computer-generated content, the first event is first computer-generated content having a first immersion level, and the first computer-generated content presented at the first immersion level occupies a first portion of the three-dimensional computer-generated environment. The system includes, in accordance with a determination that it responds to an individual request to present a first visual content within a first part of a three-dimensional environment and outputting the first audio content using a first audio output mode, and, in accordance with a determination that it responds to an individual request to present a first computer-generated content having a second immersion level, wherein the first computer-generated content presented at the second immersion level occupies a second part of the three-dimensional computer-generated environment that is larger than the first part of the three-dimensional environment, the system includes, in accordance with a determination that it responds to an individual request to present a first computer-generated content having a second immersion level, wherein the first computer-generated content presented at the second immersion level occupies a second part of the three-dimensional computer-generated environment that is larger than the first part of the three-dimensional environment, the system includes, in which using the second audio output mode instead of the first audio output mode changes the immersion level of the first audio content.
[0007] According to some embodiments, the method is performed by a computer system that communicates with a display generation component, and via the display generation component, displays a view of a computer-generated environment; detects a first movement of a first physical object in the physical environment while the computer-generated environment is being displayed and the computer-generated environment does not include a visual representation of a first part of a first physical object that is located in the physical environment where the user is located; and, in response to the detection of the first movement of the first physical object in the physical environment, determines that the user is within a threshold distance of a first part of the first physical object and that the first physical object satisfies preset criteria including requirements for preset properties of the first physical object other than the distance of the first physical object from the user, and then displays a second part of the first physical object. The present invention includes, where both the first part of the first physical object and the second part of the physical object change the appearance of a portion of virtual content displayed at a position corresponding to the current location of the first part of the first physical object without changing the appearance of a portion of virtual content displayed at a position corresponding to the second part, which is part of the range of the first physical object that is potentially visible to the user, based on the user's field of view of the computer-generated environment; and ceasing to change the appearance of a portion of virtual content displayed at a position corresponding to the current location of the first part of the first physical object, in accordance with the determination that the user is within a threshold distance of the first physical object in the physical environment surrounding the user and the first physical object does not meet a predetermined criterion.
[0008] According to some embodiments, the method is performed on a computer system that communicates with a first display generating component and one or more input devices, and includes: displaying a three-dimensional environment including a representation of a physical environment via the first display generating component; detecting a user's hand touching a specific part of the physical environment while the three-dimensional environment including the representation of the physical environment is being displayed; displaying a first visual effect at a location in the three-dimensional environment corresponding to a first part of the physical environment identified based on a scan of the first part of the physical environment, in response to the detection that the user's hand is touching a specific part of the physical environment, according to the determination that the user's hand is touching a first part of the physical environment; and displaying a second visual effect at a location in the three-dimensional environment corresponding to a second part of the physical environment identified based on a scan of the second part of the physical environment, according to the determination that the user's hand is touching a second part of the physical environment different from the first part of the physical environment.
[0009] According to some embodiments, the method is performed on a computer system that communicates with a first display generation component and one or more input devices, and via the first display generation component, displays a view of a three-dimensional environment which simultaneously includes a representation of a first part of a physical environment, including a first physical surface, and first virtual content, wherein the first virtual content includes first user interface objects displayed at positions in a three-dimensional environment corresponding to the locations of the first physical surface within the first part of the physical environment; while displaying the view of the three-dimensional environment, detect a part of the user at a first location within the first part of the physical environment, which is between the first physical surface and a viewpoint corresponding to the view of the three-dimensional environment; and in response to the detection of the part of the user at the first location within the first part of the physical environment, the representation of the part of the user is first - Includes stopping the display of the first part of the first user interface object while maintaining the display of the second part of the first user interface object so that the first part of the interface object appears to be in a previously displayed position; detecting the movement of the user part from a first location in the first part of the physical environment to a second location which lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment while displaying a view of the three-dimensional environment; and, in response to the detection of the movement of the user part from the first location to the second location, restoring the display of the first part of the first user interface object and stopping the display of the second part of the first user interface object so that the representation of the user part appears to be in a previously displayed position for the second part of the first user interface object.
[0010] According to some embodiments, the computer system includes, or communicates with, a display generating component (e.g., a display, a projector, a head-mounted display), one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, one or more processors, and one or more memory for storing or communicating with one or more programs, wherein one or more programs are configured to be executed by one or more processors, and one or more programs include instructions for performing or causing to perform any of the operations described herein. According to some embodiments, a non-temporary computer-readable storage medium has instructions stored therein, and when executed by a computer system having a display generating component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), and optionally one or more tactile output generators, the instructions cause the device to perform or cause to perform any of the operations described herein. According to some embodiments, a graphical user interface of a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, one or more sensors that optionally detect the intensity of contact with the touch-sensitive surface), one or more tactile output generators that optionally detect one or more tactile output generators, memory, and one or more processors that execute one or more programs stored in memory, includes one or more elements that are displayed in any of the methods described herein, and these elements are updated in response to input as described in any of the methods described herein. According to some embodiments, a computer system includes a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, one or more sensors that optionally detect the intensity of contact with the touch-sensitive surface), one or more tactile output generators that optionally detect one or more tactile output generators, and means for performing or causing to perform any of the operations described herein.According to some embodiments, an information processing device for use in a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensing surface, optionally one or more sensors for detecting the intensity of contact with the touch-sensing surface), and optionally one or more tactile output generators, includes means for performing or causing to perform any operation of the methods described herein.
[0011] Therefore, computer systems having display generation components are provided with improved methods and interfaces that enhance the effectiveness, efficiency, and user safety and satisfaction of such computer systems by interacting with a three-dimensional environment and simplifying the user's use of the computer system when interacting with a three-dimensional environment. Such methods and interfaces may complement or replace conventional methods for interacting with a three-dimensional environment and facilitating the user's use of the computer system when interacting with a three-dimensional environment.
[0012] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The functions and advantages described herein are not exhaustive, and many additional functions and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention. [Brief explanation of the drawing]
[0013] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.
[0014] [Figure 1] This block diagram shows the operating environment of a computer system for providing a CGR experience in several embodiments.
[0015] [Figure 2] Block diagram showing a controller for a computer system configured to manage and adjust the user's CGR experience, according to several embodiments.
[0016] [Figure 3] This block diagram shows display generation components of a computer system configured to provide the user with visual components of a CGR experience, according to several embodiments.
[0017] [Figure 4] This is a block diagram showing a hand tracking unit for a computer system configured to capture user gesture input, according to several embodiments.
[0018] [Figure 5] This is a block diagram showing an eye-tracking unit for a computer system configured to capture user eye-gaze input, according to several embodiments.
[0019] [Figure 6] This is a flowchart showing a Glint-assisted eye-tracking pipeline according to several embodiments.
[0020] [Figure 7A] This block diagram shows several embodiments of selecting different audio output modes according to the level of immersion to which computer-generated content is presented. [Figure 7B] This block diagram shows several embodiments of selecting different audio output modes according to the level of immersion to which computer-generated content is presented.
[0021] [Figure 7C]A block diagram showing the change in the appearance of a part of virtual content (e.g., enabling a part of the physical object to break through the virtual content, changing one or more visual characteristics of the virtual content based on the visual characteristics of a part of the physical object, etc.) when an important physical object approaches the display generation component or the user's location according to some embodiments. [Figure 7D] A block diagram showing the change in the appearance of a part of virtual content (e.g., enabling a part of the physical object to break through the virtual content, changing one or more visual characteristics of the virtual content based on the visual characteristics of a part of the physical object, etc.) when an important physical object approaches the display generation component or the user's location according to some embodiments. [Figure 7E] A block diagram showing the change in the appearance of a part of virtual content (e.g., enabling a part of the physical object to break through the virtual content, changing one or more visual characteristics of the virtual content based on the visual characteristics of a part of the physical object, etc.) when an important physical object approaches the display generation component or the user's location according to some embodiments. [Figure 7F] A block diagram showing the change in the appearance of a part of virtual content (e.g., enabling a part of the physical object to break through the virtual content, changing one or more visual characteristics of the virtual content based on the visual characteristics of a part of the physical object, etc.) when an important physical object approaches the display generation component or the user's location according to some embodiments. [Figure 7G] A block diagram showing the change in the appearance of a part of virtual content (e.g., enabling a part of the physical object to break through the virtual content, changing one or more visual characteristics of the virtual content based on the visual characteristics of a part of the physical object, etc.) when an important physical object approaches the display generation component or the user's location according to some embodiments. [Figure 7H]This block diagram shows some embodiments of how the appearance of parts of virtual content may change when important physical objects approach the display generation components or the user's location (for example, allowing a representation of part of the physical object to break through the virtual content, or changing one or more visual properties of the virtual content based on the visual properties of part of the physical object).
[0022] [Figure 7I] This block diagram illustrates the application of visual effects to regions within a three-dimensional environment that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment (e.g., characterized by shape, plane, and / or surface), according to several embodiments. [Figure 7J] This block diagram illustrates the application of visual effects to regions within a three-dimensional environment that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment (e.g., characterized by shape, plane, and / or surface), according to several embodiments. [Figure 7K] This block diagram illustrates the application of visual effects to regions within a three-dimensional environment that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment (e.g., characterized by shape, plane, and / or surface), according to several embodiments. [Figure 7L] This block diagram illustrates the application of visual effects to regions within a three-dimensional environment that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment (e.g., characterized by shape, plane, and / or surface), according to several embodiments. [Figure 7M] This block diagram illustrates the application of visual effects to regions within a three-dimensional environment that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment (e.g., characterized by shape, plane, and / or surface), according to several embodiments. [Figure 7N]This block diagram illustrates the application of visual effects to regions within a three-dimensional environment that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment (e.g., characterized by shape, plane, and / or surface), according to several embodiments.
[0023] [Figure 7O] This block diagram shows several embodiments of displaying an interactive user interface object at a position in a three-dimensional environment corresponding to a first part of the physical environment (e.g., a location on a physical surface, or a location in free space within the physical environment), and selectively deselecting the display of individual sub-parts of the user interface object according to the location of a part of the user (e.g., a user's finger, hand, etc.) moving in space between the first part of the physical environment and the location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment. [Figure 7P] This block diagram shows several embodiments of displaying an interactive user interface object at a position in a three-dimensional environment corresponding to a first part of the physical environment (e.g., a location on a physical surface, or a location in free space within the physical environment), and selectively deselecting the display of individual sub-parts of the user interface object according to the location of a part of the user (e.g., a user's finger, hand, etc.) moving in space between the first part of the physical environment and the location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment. [Figure 7Q] This block diagram shows several embodiments of displaying an interactive user interface object at a position in a three-dimensional environment corresponding to a first part of the physical environment (e.g., a location on a physical surface, or a location in free space within the physical environment), and selectively deselecting the display of individual sub-parts of the user interface object according to the location of a part of the user (e.g., a user's finger, hand, etc.) moving in space between the first part of the physical environment and the location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment.
[0024] [Figure 8] This is a flowchart illustrating a method for selecting different audio output modes depending on the level of immersion to which computer-generated content is presented, according to several embodiments.
[0025] [Figure 9] This is a flowchart illustrating, in several embodiments, a method for changing the appearance of a portion of virtual content when a significant physical object approaches a display generation component or the user's location.
[0026] [Figure 10] This is a flowchart of a method for applying a visual effect to a region in a three-dimensional environment that corresponds to a part of a physical environment identified based on a scan of that part of the physical environment, according to several embodiments.
[0027] [Figure 11] This is a flowchart of a method, according to several embodiments, for displaying an interactive user interface object at a position in a three-dimensional environment corresponding to a first part of a physical environment, and selectively deselecting the display of individual sub-parts of the user interface object according to the location of a part of the user moving in space between the first part of the physical environment and a location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment. [Modes for carrying out the invention]
[0028] This disclosure relates to user interfaces that provide a computer-generated reality (CGR) experience to a user, in several embodiments.
[0029] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.
[0030] In some embodiments, the computer system displays computer-generated content such as movies, virtual offices, application environments, games, and computer-generated experiences (e.g., virtual reality experiences, augmented reality experiences, or mixed reality experiences). In some embodiments, the computer-generated content is displayed in a three-dimensional environment. In some embodiments, the computer system can display visual components of computer-generated content with multiple levels of immersion, corresponding to varying degrees of emphasis on visual input from the virtual content compared to visual input from the physical environment. In some embodiments, higher levels of immersion correspond to greater emphasis on visual input from the virtual content than from the physical environment. Similarly, in some embodiments, audio components of computer-generated content accompanying and / or corresponding to the visual components of the computer-generated content (e.g., sound effects and soundtracks in movies; audio alerts, voice feedback, and system sounds in application environments; sound effects, speech, and voice feedback in games; and / or sound effects and voice feedback in computer-generated experiences) can be output at multiple levels of immersion. In some embodiments, multiple levels of immersion optionally correspond to varying degrees of spatial correspondence between the position of a virtual sound source in virtual content displayed via a display-generated element and the perceived location of the virtual sound source within a selected reference frame of the virtual sound source. In some embodiments, the selected reference frame of an individual virtual sound source is based on the physical environment, the virtual three-dimensional environment of the computer-generated content, the viewpoint of the currently displayed view of the three-dimensional environment of the computer-generated content, the location of the display-generated element in the physical environment, or the user's location in the physical environment.In some embodiments, a higher level of immersion corresponds to a higher level of correspondence between the position of the virtual sound source in the computer-generated environment and the perceived location of the virtual sound source in a selected reference frame of the audio component of the computer-generated content (e.g., a reference frame based on the three-dimensional environment depicted in the computer-generated experience, a reference frame based on the viewpoint location, a reference frame based on the location of the display-generated element, a reference frame based on the user's location, etc.). In some embodiments, a lower level of correspondence between the position of the virtual sound source in the computer-generated environment and the perceived location of the sound source in a selected reference frame of the audio component of the computer-generated content is a result of a higher level of correspondence between the perceived location of the virtual sound source and the location of the audio output device in the physical environment (e.g., the sound appears to originate from the location of the audio output device, regardless of the position of the virtual sound source in the three-dimensional environment of the computer-generated content, and / or regardless of the viewpoint location, the location of the display-generated element, and / or the user's location, etc.). In some embodiments, the computer system detects a first event corresponding to a request to present a first computer-generated experience, and the computer system selects an audio output mode for outputting the audio components of the computer-generated experience according to the level of immersion at which the visual components of the computer-generated experience are displayed via the display-generated components. At a higher level of immersion associated with the display of the visual content of the first computer-generated experience, the computer system selects an audio output mode for presenting the audio content of the computer-generated experience having a corresponding higher level of immersion. In some embodiments, displaying visual content having a higher level of immersion includes displaying the visual content in a larger spatial range in a three-dimensional environment, and outputting audio content having a corresponding higher level of immersion includes outputting the audio content in a spatial audio output mode.In some embodiments, when switching the display of visual content having two different levels of immersion (e.g., from a higher level of immersion to a lower level of immersion, or from a lower level of immersion to a higher level of immersion), the computer system also switches the output of audio content having two different levels of immersion (e.g., from spatial audio output mode to stereo audio output mode, from surround sound output mode to stereo audio output mode, from stereo audio output mode to surround sound output mode, or from stereo audio output mode to spatial audio output mode). Selecting an appropriate audio output mode for outputting the audio component of computer-generated content according to the level of immersion at which the visual content of the computer-generated content is displayed allows the computer system to provide a computer-generated experience that better matches user expectations and avoids causing confusion when the user interacts with the computer-generated environment while engaging with the computer-generated experience. This can reduce user errors and make user interaction with the computer system more efficient.
[0031] In some embodiments, when virtual content is displayed in a three-dimensional environment (e.g., a virtual reality environment, an augmented reality environment, etc.), all or part of the view of the physical environment is blocked or replaced by the virtual content. In some cases, it is advantageous to give display priority to certain physical objects in the physical environment over virtual content so that at least a portion of the physical objects are visually represented in the view of the three-dimensional environment. In some embodiments, the computer system utilizes various criteria for determining whether to give display priority to individual physical objects so that the representation of the individual physical object can overwrite the portion of the virtual content currently displayed in the three-dimensional environment, when the location of the individual physical object in the physical environment corresponds to the position of a portion of the virtual content in the three-dimensional environment. In some embodiments, the criteria include the requirement that at least a portion of the physical object has entered an approaching threshold space region surrounding the user of the display-generating component (e.g., a user viewing virtual content through the display-generating component, a user whose view of a portion of the physical object is blocked or replaced by the display of the virtual content, etc.), and the additional requirement that the computer system detects the presence of one or more characteristics of the physical object that indicate to the user the increased importance of the physical object. In some embodiments, a physical object of increased importance to the user may be a friend or family member, a team member or supervisor, or a pet. In some embodiments, a physical object of increased importance to the user may be a person or object that requires the user's attention to deal with an emergency. In some embodiments, a physical object of greater importance to the user may be a person or object that requires the user's attention to take action that the user does not miss. The criteria can be adjusted by the user based on the user's needs and desires, and / or by the system based on contextual information (e.g., time, location, scheduled event, etc.).In some embodiments, giving display priority to physical objects that are more important than virtual content, and visually representing at least a portion of physical objects in a view of the three-dimensional environment, includes replacing the display of a portion of virtual content with a representation of a portion of physical objects, or changing the appearance of a portion of virtual content according to the appearance of a portion of physical objects. In some embodiments, at least a portion of a physical object remains not visually represented in the view of the three-dimensional environment, and remains blocked or replaced by the display of virtual content, even if the position corresponding to the location of the portion of the physical object is visible within the field of view provided by the display generation component (e.g., the position is currently occupied by virtual content). In some embodiments, portions of the three-dimensional environment that have been modified to indicate the presence of physical objects, and portions of the three-dimensional environment that have not been modified to indicate the presence of physical objects (e.g., portions of the three-dimensional environment may continue to change based on the progress of the computer-generated experience and / or user interaction with the three-dimensional environment, etc.), correspond to positions on a continuum of virtual objects or surfaces. The system allows at least a portion of important physical objects to break the display of virtual content and be visually represented in a position corresponding to the location of some of the physical objects, while at least a portion of the physical objects remain visually hidden by the virtual content. The system provides the user with the opportunity to perceive and interact with physical objects without completely interrupting the computer-generated experience in which the user is involved, and without indiscriminately allowing physical objects of little importance to the user (e.g., a rolling ball, a passerby) to interrupt the computer-generated experience, based on the determination that the physical objects meet predefined criteria for identifying physical objects of increased importance to the user and that the physical objects have entered a predefined spatial area surrounding the user.This improves the user experience and reduces the number, extent, and / or nature of user input to achieve desired results (e.g., manually stopping the computer-generated experience when physically interrupted or touching a physical object, or manually restarting the computer-generated experience after an unnecessary interruption), thereby creating a more efficient human-machine interface.
[0032] In some embodiments, the computer system displays a representation of the physical environment in response to a request to display a three-dimensional environment, including a representation of the physical environment (e.g., in response to a user putting on a head-mounted display, in response to a user request to start an augmented reality environment, in response to a user request to end a virtual reality experience, in response to a user turning on or waking up display generation components from a low-power state, etc.). In some embodiments, the computer system initiates a scan of the physical environment to identify objects and surfaces within the physical environment and optionally constructs a three-dimensional or pseudo-three-dimensional model of the physical environment based on the identified objects and surfaces within the physical environment. In some embodiments, the computer system initiates a scan of the physical environment in response to receiving a request to display the three-dimensional environment (e.g., if the physical environment has not been previously scanned and characterized by the computer system, or if a rescan is requested by the user or the system based on a preset rescan criterion that is met (e.g., the last scan was performed more than a threshold time ago, the physical environment has changed, etc.)). In some embodiments, the computer system initiates a scan in response to detecting a user's hand touching a part of the physical environment (e.g., a physical surface, a physical object, etc.). In some embodiments, the computer system initiates a scan upon detection that the user's line of sight, directed to a position corresponding to a portion of the physical environment, meets pre-set stability and / or duration criteria. In some embodiments, the computer system displays visual feedback regarding the progress and results of the scan (e.g., identification of physical objects and surfaces in the physical environment, determination of the physical and spatial properties of physical objects and surfaces). In some embodiments, the visual feedback includes displaying individual visual effects on distinct parts of the three-dimensional environment that are touched by the user's hand and correspond to the portion of the physical environment identified based on the scan of that portion of the physical environment. In some embodiments, the visual effects include representations of motion extending from and / or propagating from the distinct parts of the three-dimensional environment.In some embodiments, the computer system displays visual effects in response to the detection of the user's hand touching individual parts of the physical environment, and the three-dimensional environment is displayed after the scan of the physical environment is complete, in response to a previous request to display the three-dimensional environment. In some embodiments, displaying visual effects indicating the progress and results of the scan of the physical environment at positions corresponding to the location of the user's touch on parts of the physical environment helps the user visualize the spatial environment that the computer uses to display and fix virtual objects and virtual surfaces, facilitating subsequent interaction between the user and the spatial environment. This makes the interaction more efficient, reduces input errors, and creates a more efficient human-machine interface. In some embodiments, the location of the user's contact with parts of the physical environment is utilized by the computer system to generate a three-dimensional model of the physical environment and provide more accurate boundary conditions for identifying surface and object boundaries based on the scan, making the display of virtual objects more accurate and seamless in the three-dimensional environment.
[0033] In some embodiments, a computer system displays interactive user interface objects in a three-dimensional environment. The computer system also displays a representation of the physical environment within the three-dimensional environment, and the interactive user interface objects have distinct spatial relationships to various positions in the three-dimensional environment corresponding to different locations in the physical environment. When a user interacts with the three-dimensional environment with a part of the user's hand, such as one or more fingers or the entire hand, through touch input and / or gesture input, the user's hand and possibly the wrist and arm connected to the hand can enter a spatial region between locations corresponding to the positions of user interface objects (e.g., locations of physical objects or physical surfaces, locations in free space within the physical environment, etc.) and locations corresponding to the viewpoint of the currently displayed view of the three-dimensional environment (e.g., the location of the user's eyes, the location of display-generating components, the location of a camera capturing a view of the physical environment shown in the three-dimensional environment, etc.). Based on the spatial relationships between the location of the user's hand, the locations corresponding to the positions of user interface objects, and the locations corresponding to the viewpoint, the computer system determines which parts of the user interface objects are visually blocked by the user's parts and which parts of the user interface objects are not visually blocked by the user's parts when viewed by the user from the viewpoint location. The computer system then stops displaying individual parts of a user interface object that would be visually blocked by the user's part (as determined by the computer system, for example), and instead allows the user's representation to be seen at the position of the individual part of the user interface object, while maintaining the display of other parts of the user interface object that would not be visually blocked by the user's part (as determined by the computer system, for example).In some embodiments, in response to the detection of movement of a user part or viewpoint (e.g., due to movement of a display generation component, movement of a camera capturing the physical environment, movement of the user's head or torso), the computer system re-evaluates which parts of the user interface object are visually blocked by the user part and which parts are not visually blocked by the user part when viewed by the user from the viewpoint location, based on the new spatial relationships between the user part, the location corresponding to the viewpoint, and the location corresponding to the position of the user interface object. The computer system then stops displaying the other parts of the user interface object that will be visually blocked by the user part (e.g., as determined by the computer system), and allows the previously stopped displaying parts of the user interface object to be restored to the view of the three-dimensional environment. Visually segmenting the user interface object into multiple parts and replacing the display of one or more parts of the user interface object with a representation of the user part that has entered the spatial region between the location corresponding to the position of the user interface object and the location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment helps the user visualize and perceive the placement location of the user interface object relative to their hand, and facilitates interaction between the user and the user interface object in the three-dimensional environment. This makes interactions more efficient, reduces input errors, and creates a more efficient human-machine interface.
[0034] Figures 1 to 6 illustrate exemplary computer systems for providing a CGR experience to a user. Figures 7A to 7B are block diagrams illustrating, in several embodiments, the selection of different audio output modes according to the level of immersion to which computer-generated content is presented. Figures 7C to 7H are block diagrams illustrating, in several embodiments, the alteration of the appearance of parts of virtual content when a significant physical object approaches a display-generated component or the user's location of a display-generated component. Figures 7I to 7N are block diagrams illustrating, in several embodiments, the application of visual effects to areas in a three-dimensional environment corresponding to a part of the physical environment identified based on a scan of that part of the physical environment. Figures 70 to 7Q are block diagrams illustrating, in several embodiments, the display of an interactive user interface object at a position in a three-dimensional environment corresponding to a first part of the physical environment, and the selective removal of the display of individual sub-parts of the user interface object according to the location of the user moving in space between the first part of the physical environment and a location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment. The user interfaces of Figures 7A to 7Q are used to illustrate each of the processes in Figures 8 to 11.
[0035] In some embodiments, as shown in Figure 1, the CGR experience is provided to the user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or remote server), display generation components 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, within a head-mounted device or handheld device).
[0036] When describing a CGR experience, various terms are used to refer individually to several related but distinct environments that the user perceives and / or interacts with (for example, using inputs detected by the computer system 101, which causes the computer system generating the CGR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the CGR experience). The following is a subset of these terms.
[0037] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.
[0038] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In CGR, a subset of a person's bodily movements or their representations are tracked, and in response, one or more properties of one or more virtual objects simulated within the CGR environment are adjusted to behave according to at least one law of physics. For example, a CGR system may detect a person's head rotation and, in response, adjust the graphic content and sound field presented to the person in a similar manner to how such views and sounds would change in a physical environment. Depending on the circumstances (e.g., for reasons of accessibility), adjustments to the properties(s) of virtual objects(s) in the CGR environment may be made in response to representations of bodily movements (e.g., voice commands). A person may perceive and / or interact with CGR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person can perceive and / or interact with an audio object that creates a 3D or spatial audio environment, providing the perception of a point sound source in 3D space. In another example, an audio object may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, a person may perceive and / or interact only with audio objects.
[0039] Examples of CGR include virtual reality and mixed reality.
[0040] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their bodily movements within the computer-generated environment.
[0041] Mixed Reality: A mixed reality (MR) environment, in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input, refers to a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects). On a virtual continuum, a mixed reality environment is any location between, but not encompassing, the complete physical environment at one end and the virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track location and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.
[0042] Examples of mixed reality include augmented reality and augmented virtual reality.
[0043] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through images or videos of the physical environment. As used herein, videos of the physical environment shown on an opaque display are referred to as “pass-through videos,” and it means that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, thereby allowing a person to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the image sensor. As another example, the representation of the physical environment may be transformed by graphically altering (e.g., enlarging) a portion of it, thereby making the altered portion a modified version that represents the original captured image but is not photorealistic. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.
[0044] Augmented Virtual: An Augmented Virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, while people with faces are realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.
[0045] Hardware: There are many different types of electronic systems that enable people to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to the human eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination thereof. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the human retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces. In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience.In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., physical setup / environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is communicably coupled to a display generation component 120 (e.g., an HMD, display, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within a housing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0046] In some embodiments, the display generation component 120 is configured to provide the user with a CGR experience (e.g., at least the visual components of the CGR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0047] According to some embodiments, the display generation component 120 provides the user with a CGR experience while the user is virtually and / or physically present in the scene 105.
[0048] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., their head or hand). Thus, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with CGR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the CGR content response is displayed via the HMD. Similarly, a user interface showing interaction with CGR content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the interaction is triggered by the movement of the HMD relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).
[0049] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features for the sake of simplification are not shown so as not to obscure more suitable embodiments of the exemplary embodiments disclosed herein.
[0050] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0051] In some embodiments, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0052] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and CGR experience module 240.
[0053] The operating system 230 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). For this purpose, in various embodiments, the CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.
[0054] In some embodiments, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 in Figure 1, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data acquisition unit 242 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0055] In some embodiments, the tracking unit 244 is configured to map scene 105 and track the position / location of at least the display generation components 120 relative to scene 105 in Figure 1, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, the tracking unit 244 includes a hand tracking unit 245 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 245 is configured to track the position / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to scene 105 in Figure 1, relative to the display generation components 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 245 is described in more detail below with respect to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to the CGR content displayed via the display generation component 120. The eye-tracking unit 243 will be described in more detail below with reference to Figure 5.
[0056] In some embodiments, the adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0057] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0058] While the data acquisition unit 242, tracking unit 244 (including, for example, the eye-tracking unit 243 and the hand-tracking unit 245), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, tracking unit 244 (including, for example, the eye-tracking unit 243 and the hand-tracking unit 245), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.
[0059] Furthermore, Figure 2 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0060] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting embodiments, the HMD120 may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0061] In some embodiments, one or more communication buses 304 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0062] In some embodiments, one or more CGR displays 312 are configured to provide the user with a CGR experience. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, the HMD 120 includes a single CGR display. In another embodiment, the HMD 120 includes a CGR display for each of the user's eyes. In some embodiments, one or more CGR displays 312 can present MR or VR content. In some embodiments, one or more CGR displays 312 can present MR or VR content.
[0063] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.
[0064] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and CGR presentation module 340.
[0065] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to the user via one or more CGR displays 312. For this purpose, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.
[0066] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. For this purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0067] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. For this purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0068] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment on which computer-generated objects can be placed) based on media content data. For this purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0069] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral devices 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0070] Although the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 are shown as existing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0071] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0072] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by a hand tracking unit 245 (Figure 2) to track the position / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to a coordinate system defined for the scene 105 in Figure 1 (e.g., relative to parts of the physical environment surrounding the user, relative to the display generation component 120, or relative to parts of the user (e.g., the user's face, eyes, or head) and / or relative to the user's hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0073] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.
[0074] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which drives the display generation components 120 accordingly. For example, a user may interact with the software running on the controller 110 by moving their hand 408 and changing the hand's orientation.
[0075] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines a set of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the hand tracking device 440 may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.
[0076] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D location of the user's wrist and fingertips.
[0077] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The posture estimation function described herein may be interleaved with the motion tracking function, so that patch-based posture estimation is performed only once every two (or more) frames, while tracking is used to detect changes in posture that occur over the remaining frames. Posture, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 in response to the posture and / or gesture information, or perform other functions.
[0078] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 440, but some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the hand tracking device 402, or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (for example, in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.
[0079] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, e.g., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values to identify and segment image components (e.g., adjacent pixel groups) that have features of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.
[0080] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.
[0081] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 243 (Figure 2) to track the position and movement of the user's gaze relative to the scene 105 or to the CGR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating CGR content for user viewing and a component for tracking the user's gaze relative to the CGR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or CGR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally part of a non-head-mounted display generation component.
[0082] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0083] As shown in Figure 5, in some embodiments, the eye-tracking device 130 includes at least one eye-tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visual light to pass through. The eye-tracking device 130 optionally captures images of the user's eye (e.g., as a video stream captured at 60 to 120 frames per second (fps)), analyzes the images to generate eye-tracking information, and communicates the eye-tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a separate eye-tracking camera and light source.
[0084] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil location, central visual location, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a glint-assisted method to determine the user's current visual axis and viewpoint relative to the display.
[0085] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece (one or more) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes (one or more) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display or projector of a handheld device) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the user's eye(s) 592 (e.g., as shown at the bottom of Figure 5).
[0086] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses eye-tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the eye-tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the eye-tracking input 542 is optionally used to determine the direction the user is currently looking.
[0087] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the CGR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can utilize the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.
[0088] In some embodiments, the eye-tracking device is part of a head-mounted device mounted on a wearable housing, which includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., a light source 530 (e.g., an IR LED or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.
[0089] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0090] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality (including, for example, virtual reality and / or mixed reality) applications to provide users with computer-generated reality (including, for example, virtual reality, augmented reality and / or augmented virtual reality) experiences.
[0091] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in a tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in a tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.
[0092] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which is initiated at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0093] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.
[0094] At 640, if the process proceeds from element 410, the current frame is analyzed and the pupil and glint are tracked, based in part on prior information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to confirm that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.
[0095] Figure 6 is intended to serve as an example of an eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in the computer system 101 to provide the user with a CGR experience in various embodiments, either in place of or in combination with the glint-assisted eye-tracking technology described herein.
[0096] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, each example may be compatible with the input device or method described in the other example, and their use should be considered optional. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, each example may be compatible with the output device or method described in the other example, and their use should be considered optional. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, each example may be compatible with the method described in the other example, and their use should be considered optional. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.
[0097] User interface and related processes Here, we focus on embodiments of a user interface ("UI") and related processes that may be performed in a computer system such as a portable multifunction device or head-mounted device, which comprises display generation components, one or more input devices, and (optionally) one or more cameras.
[0098] Figures 7A to 7Q illustrate three-dimensional environments displayed via display generation components (e.g., display generation component 7100, display generation component 120, etc.) according to various embodiments, and interactions occurring in the three-dimensional environment caused by user input to the three-dimensional environment. In some embodiments, input is directed to a virtual object in the three-dimensional environment by the user's gaze detected at the position of the virtual object, a hand gesture performed at a location in the physical environment corresponding to the position of the virtual object, or a hand gesture performed at a location in the physical environment independent of the position of the virtual object while the virtual object has input focus (e.g., selected by gaze, selected by pointer, selected by previous gesture input, etc.). In some embodiments, input is directed to a representation of a physical object or a virtual object corresponding to a physical object by the user's hand movements (e.g., movement of the entire hand, movement of the entire hand in individual poses, movement of one part of the hand relative to another part of the hand, relative movement between two hands, etc.) and / or operations on a physical object (e.g., touch, swipe, tap, open, move toward, move relatively, etc.).
[0099] In some embodiments, the three-dimensional environment displayed via the display generation component is a virtual three-dimensional environment containing virtual objects and content in different virtual positions within the three-dimensional environment without a representation of the physical environment. In some embodiments, the three-dimensional environment is a mixed reality environment that displays virtual objects in different virtual positions within the three-dimensional environment constrained by one or more physical aspects of the physical environment (e.g., the position and orientation of walls, floors, and surfaces, the direction of gravity, time, etc.). In some embodiments, the three-dimensional environment is an augmented reality environment that includes a representation of the physical environment. The representation of the physical environment includes respective representations of physical objects and surfaces at different positions within the three-dimensional environment, such that the spatial relationships between different physical objects and surfaces within the physical environment are reflected by the spatial relationships between the representations of physical objects and surfaces within the three-dimensional environment. When virtual objects are positioned relative to the positions of the representations of physical objects and surfaces within the three-dimensional environment, they appear to have corresponding spatial relationships with the physical objects and surfaces within the physical environment.
[0100] In some embodiments, the display generation component includes a pass-through portion on which a representation of the physical environment is displayed. In some embodiments, the pass-through portion is a transparent or translucent (e.g., see-through) portion of the display generation component that surrounds the user's field of view and reveals at least a portion of the physical environment within the field of view. For example, the pass-through portion is a translucent (e.g., less than 50%, 40%, 30%, 20%, 15%, 10%, or 5% opacity) or transparent portion of a head-mounted display or head-up display, thereby allowing the user to see the real world surrounding them through it without removing the head-mounted display or moving away from the head-up display. In some embodiments, the pass-through portion gradually transitions from translucent or transparent to completely opaque when displaying a virtual or mixed reality environment. In some embodiments, the pass-through portion of the display generation component displays a live feed of images or videos of at least a portion of the physical environment captured by one or more cameras (e.g., one or more rear cameras associated with a mobile device or head-mounted display, or other cameras supplying image data to an electronic device). In some embodiments, one or more cameras are directed towards a part of the physical environment that is directly in front of the user (e.g., behind the display generating components). In some embodiments, one or more cameras are directed towards a part of the physical environment that is not directly in front of the user (e.g., in a different physical environment, or to the side or behind the user).
[0101] In some embodiments, when displaying virtual objects at positions corresponding to the locations of one or more physical objects in a physical environment, at least some of the virtual objects are displayed in a configuration of a portion of the camera's live view (e.g., a portion of the physical environment captured in the live view) (e.g., replacing that display). In some embodiments, at least some of the virtual objects and content are projected onto a physical surface or empty space in the physical environment and are visible through pass-through portions of the display-generating component (e.g., as part of the camera view of the physical environment, or through transparent or translucent portions of the display-generating component, etc.). In some embodiments, at least some of the virtual objects and content are displayed so as to overlay a portion of the display, blocking, but not all, the view of at least some of the physical environment visible through transparent or translucent portions of the display-generating component. In some embodiments, at least some of the virtual objects are projected directly onto the user's retina in a position relative to an image representation of the physical environment (e.g., visible via the camera view of the physical environment, or through transparent portions of the display-generating component, etc.).
[0102] In some embodiments, the display generation component displays different views of the three-dimensional environment in accordance with user input or movement that changes the virtual position of the viewpoint of the currently displayed view of the three-dimensional environment relative to the three-dimensional environment. In some embodiments, if the three-dimensional environment is a virtual environment, the viewpoint moves in accordance with navigation or movement requests (e.g., aerial hand gestures, gestures performed by the movement of one part of the hand relative to another part of the hand) without requiring movement of the user's head, torso, and / or the display generation component in the physical environment. In some embodiments, movement of the user's head and / or torso relative to the physical environment, and / or movement of the display generation component of the computer system or other location sensing elements (e.g., due to the user holding the display generation component or wearing an HMD) causes corresponding movement of the viewpoint relative to the three-dimensional environment (e.g., corresponding changes in direction, distance, speed, and / or orientation), resulting in a corresponding change in the currently displayed view of the three-dimensional environment. In some embodiments, if a virtual object has a predefined spatial relationship with respect to a viewpoint, movement of the viewpoint relative to the three-dimensional environment causes movement of the virtual object relative to the three-dimensional environment while the position of the virtual object in the field of view is maintained (for example, the virtual object is said to be headlocked). In some embodiments, the virtual object is bodylocked to the user and moves relative to the three-dimensional environment when the user moves as a whole in the physical environment (e.g., carrying or wearing a computer system display generation component and / or other location sensing component), but does not move in the three-dimensional environment in response to the movement of the user's head (e.g., a computer system display generation component and / or other location sensing component rotating around the user's fixed location in the physical environment).
[0103] In some embodiments, the view of the three-dimensional environment shown in Figures 7A to 7Q includes representations of the user's hand(s), arm(s), and / or wrist(s). In some embodiments, the representations are part of the representation of the physical environment provided via a display generation component. In some embodiments, the representations are not part of the representation of the physical environment but are captured separately (e.g., by one or more cameras pointed at the user's hand(s), arm(s), and wrist(s)) and displayed in the three-dimensional environment independently of the view of the three-dimensional environment. In some embodiments, the representations include camera images captured by one or more cameras of a computer system(s), or stylized versions of the arm, wrist, and / or hand based on information captured by various sensors. In some embodiments, the representations replace a representation of part of the representation of the physical environment, overlay on part of the representation of the physical environment, or block a view of part of the representation of the physical environment. In some embodiments, when the display generation component does not provide a view of the physical environment and provides a completely virtual environment (e.g., without a camera field of view or transparent pass-through portion), a real-time visual representation of one or both of the user's arms, wrists, and / or hands (e.g., a stylized representation or a segmented camera image) may still be displayed in the virtual environment. In some embodiments, a representation of the user's hands is shown in the figures, but it should be understood that, unless clarified by the corresponding description, a representation of the user's hands is not necessarily always displayed and / or may not need to be displayed within the user's field of view when providing the input necessary to interact with the three-dimensional environment.
[0104] Figures 7A and 7B are block diagrams illustrating the selection of different audio output modes according to the level of immersion to which computer-generated content is presented, according to several embodiments.
[0105] In some embodiments, the computer system displays computer-generated content such as movies, virtual offices, application environments, games, and computer-generated experiences (e.g., virtual reality experiences, augmented reality experiences, mixed reality experiences, etc.). In some embodiments, the computer-generated content is displayed in a three-dimensional environment (e.g., environment 7102 in Figures 7A-7B, or another environment). In some embodiments, the computer system can display the visual components of computer-generated content (e.g., visual content 7106, or other visual content) with multiple levels of immersion, corresponding to varying degrees of emphasis on visual input from the virtual content compared to visual input from the physical environment. In some embodiments, higher levels of immersion correspond to greater emphasis on visual input from the virtual content than from the physical environment. Similarly, in some embodiments, the audio components of computer-generated content (e.g., sound effects and soundtracks in movies; audio alerts, voice feedback, and system sounds in application environments; sound effects, speech, and voice feedback in games; and / or sound effects and voice feedback in computer-generated experiences, etc.) accompanying and / or corresponding to the visual components of the computer-generated content can be output at multiple levels of immersion. In some embodiments, multiple levels of immersion optionally correspond to varying degrees of spatial correspondence between the position of a virtual sound source within virtual content displayed via a display-generated element and the perceived location of the virtual sound source within a selected reference frame of the virtual sound source. In some embodiments, the selected reference frame of an individual virtual sound source is based on the physical environment, the virtual three-dimensional environment of the computer-generated content, the viewpoint of the currently displayed view of the three-dimensional environment of the computer-generated content, the location of the display-generated element in the physical environment, or the user's location in the physical environment.In some embodiments, a higher level of immersion corresponds to a higher level of correspondence between the position of the virtual sound source in the computer-generated environment and the perceived location of the virtual sound source in a selected reference frame of the audio component of the computer-generated content (e.g., a reference frame based on the three-dimensional environment depicted in the computer-generated experience, a reference frame based on the viewpoint location, a reference frame based on the location of the display-generated element, a reference frame based on the user's location, etc.). In some embodiments, a lower level of correspondence between the position of the virtual sound source in the computer-generated environment and the perceived location of the sound source in the selected reference frame of the audio component of the computer-generated content is a result of a higher level of correspondence between the perceived location of the virtual sound source and the location of the audio output device in the physical environment (e.g., the sound appears to originate from the location of the audio output device regardless of the position of the virtual sound source in the three-dimensional environment of the computer-generated content, and / or regardless of the viewpoint location, the location of the display-generated element, and / or the user's location, etc.). In some embodiments, the computer system detects a first event corresponding to a request that presents a first computer-generated experience (e.g., requests 7112, 7114, etc. in Figures 7A-7B, or other requests, etc.), and the computer system selects an audio output mode for outputting the audio components of the computer-generated experience, depending on the level of immersion at which the visual components of the computer-generated experience are displayed via the display generation component. At a higher level of immersion associated with the display of the visual content of the first computer-generated experience, the computer system selects an audio output mode for presenting the audio content of the computer-generated experience having a correspondingly higher level of immersion. In some embodiments, displaying visual content having a higher level of immersion includes displaying the visual content in a larger spatial range in a three-dimensional environment (e.g., as shown in Figure 7B, in contrast to Figure 7A), and outputting audio content having a corresponding higher level of immersion includes outputting the audio content in a spatial audio output mode.In some embodiments, when switching the display of visual content having two different levels of immersion (e.g., from a higher level of immersion to a lower level of immersion, or from a lower level of immersion to a higher level of immersion), the computer system also switches the output of audio content having two different levels of immersion (e.g., from spatial audio output mode to stereo audio output mode, from surround sound output mode to stereo audio output mode, from stereo audio output mode to surround sound output mode, or from stereo audio output mode to spatial audio output mode).
[0106] As described herein, audio output devices, including standalone speakers (e.g., soundbars, external speakers), built-in audio output components of displays or computer systems (e.g., built-in speakers in head-mounted display devices, touchscreen display devices, portable electronic devices, or head-up displays), and wearable audio output devices (e.g., headphones, earphones, earcups, and earbuds), are widely used to provide audio output to users. The same audio content may have different audio characteristics that make the audio content sound different to the user perceiving the audio output when output using different audio output devices and / or different output modes of the same audio output device. For this reason, it is desirable to adjust the audio output mode, including changing the audio characteristics, sound source characteristics, and / or audio output device characteristics, based on the level of immersion provided to the user by the visual content of the computer-generated experience, so that when the computer-generated experience is provided to the user, the audio and visual content of the computer-generated experience harmonize and complement each other more seamlessly.
[0107] Existing stereo and mono audio output modes provide audio relative to a reference frame associated with the audio output device. In the case of a fixed audio output device, the sound appears to originate from the location of the audio output device in the physical environment, regardless of the user's movement in the physical environment and changes in the visual content of the computer-generated experience (e.g., changes due to the movement of a virtual sound source and / or viewpoint in the three-dimensional environment of the computer-generated experience). In the case of a wearable audio output device that remains stationary relative to a part of the user's body (e.g., ear, head), the sound appears to be locked to a part of the user's body, regardless of changes in the visual content of the computer-generated experience in the three-dimensional environment of the computer-generated experience (e.g., changes due to the movement of a virtual sound source, changes due to viewpoint movement (e.g., viewpoint movement triggered by a movement request from the user or computer system, and not caused by, and not corresponding to, movement of a part of the user's body)). In some cases, the audio output device and display generation components of a computer system may be housed separately and move relative to each other in the physical environment during the presentation of computer-generated content via the audio output device and display generation components. In such cases, regardless of the location of the display generation components in the physical environment, or changes in the visual content of the computer-generated experience (e.g., changes due to the movement of virtual sound sources and / or viewpoint movements in the three-dimensional environment of the computer-generated experience (e.g., movements triggered by movement requests or in response to the movement of the user or a part thereof in the physical environment)), the sound still appears to be originating from the audio output device. Thus, stereo and mono audio output modes provide a less immersive listening experience and unrealistic sounds than spatial audio output modes when the audio content of the computer-generated experience is provided to the user using stereo or mono audio output modes.
[0108] In some embodiments, the spatial audio output mode simulates a more realistic listening experience in which the audio appears to be coming from a sound source in a separate reference frame, such as a three-dimensional environment displayed via a display generation component (e.g., an augmented reality environment, a virtual reality environment, or a pure pass-through view of the physical environment surrounding the user), and the position of the simulated sound source is separated from the location and movement of the audio output device in the physical environment.
[0109] In some embodiments, the reference frame for the spatial audio output mode is based on the physical environment represented in the three-dimensional environment of the computer-generated experience, and the reference frame is optionally not affected by the user's movement, the movement of the audio output device, and / or the movement of the display generation components in the physical environment.
[0110] In some embodiments, the reference frame for the spatial audio output mode is based on the virtual three-dimensional environment of the computer-generated experience. In some embodiments, if these movements do not cause corresponding movements in the virtual three-dimensional environment, the reference frame is optionally not changed due to user movements in the physical environment, movements of the audio output device, and / or movements of the display generation components.
[0111] In some embodiments, the reference frame of the spatial audio output mode is based on the three-dimensional environment associated with the viewpoint of the currently displayed view of the three-dimensional environment. In some embodiments, if these movements do not cause a corresponding movement of the viewpoint of the currently displayed view of the three-dimensional environment, the reference frame is optionally not changed due to user movement in the physical environment, movement of the audio output device, and / or movement of display generation components.
[0112] In some embodiments, the reference frame for audio content output in spatial audio mode is optionally different from the reference frame for visual content in computer-generated experiences. For example, in some embodiments, visual content is displayed against a reference frame associated with a physical or virtual environment visually presented via display generation components, while at least a portion of virtual sound sources (e.g., external narrator, internal dialogue, etc.) are within a reference frame associated with the user's viewpoint.
[0113] In some embodiments, the audio content of the computer-generated experience optionally includes sound sources linked to different reference frames, such as a first reference frame of virtual sound sources that do not have a corresponding virtual position within the three-dimensional environment of the computer-generated experience (e.g., system-level sounds, external narration, etc.), a second reference frame of virtual sound sources that have a corresponding visual embodiment within the three-dimensional environment of the computer-generated experience (e.g., virtual objects, virtual surfaces, virtual light, etc.), and a third reference frame of virtual sound sources that are far from the viewpoint, outside the field of view, or hidden (e.g., ambient noise such as waves, insects, wind, rain, jungle sounds, etc.). In some embodiments, the first reference frame is fixed to the user's head, display-generated elements, and / or viewpoint and optionally moves. In some embodiments, the second reference frame is tied to the three-dimensional environment of the computer-generated experience and optionally moves with the display-generated elements. In some embodiments, the third reference frame is tied to the physical environment and optionally does not move with the user, display-generated elements, or viewpoint. In relation to providing visual content using display generation components, the computer system can select and configure spatial audio modes to provide a more realistic and immersive listening experience, and output sound based on the visual content presented via the display generation component, based on the spatial configuration between the audio output device and the display generation component in the physical environment, and based on the spatial configuration between the user, the display generation component, and the audio output device.
[0114] In some embodiments, the spatial speech output mode is a mode that allows speech output from a speech output device(s) to sound as if the speech were coming from one or more locations (e.g., one or more sound sources) within a selected individual reference frame for virtual sound sources such as a three-dimensional or physical environment of a computer-generated experience, where the positioning of the one or more simulated or perceived sound sources is separated from or independent of the movement of the speech output device(s)(s) relative to the individual reference frame. Typically, when one or more perceived sound sources are fixed, they are fixed relative to the individual reference frame associated with the sound sources, and when they move, they move relative to the individual reference frame.
[0115] In some embodiments, the reference frame is a reference frame based on the physical environment represented in a computer-generated experience provided through the display generation components of a computer system. In some embodiments, when the reference frame is based on the physical environment (e.g., the computer-generated experience is an augmented reality experience based on the physical environment, or a pass-through view of the physical environment), one or more perceived sound sources have their respective spatial locations within the physical environment. For example, in some embodiments, the computer-generated experience includes visual counterparts of the perceived sound sources (e.g., virtual objects that produced the sound in the computer-generated experience) having their respective positions corresponding to their respective spatial locations within the physical environment. In some embodiments, the computer-generated experience includes sound that does not include visual counterparts (e.g., remote or hidden virtual objects that produced the sound in the computer-generated experience, virtual wind, sound effects, external narrators, etc.) but has an origin corresponding to its respective spatial location within the physical environment. In some embodiments, as one or more audio output devices move around the physical environment, the audio output from one or more audio output devices is adjusted to continue emitting sound as if the sound were coming from one or more perceived sound sources located at their respective spatial locations within the physical environment. If one or more perceived sound sources are moving sound sources that travel through a set of spatial locations around the physical environment, the audio output from the audio output device(s) is adjusted to continue emitting sound as if the sound were coming from one or more perceived sound sources in a set of spatial locations within the physical environment. Such adjustments for moving the sound sources also take into account any movement of the audio output device(s) relative to the physical environment (for example, if the audio output device(s) move relative to the physical environment along a similar path as a moving sound source to maintain a constant spatial relationship with the moving sound source, the audio is output in such a way that the sound does not appear to be moving relative to the audio output device(s)).In some embodiments, when audio content is output using a spatial audio output mode and reference frame based on the physical environment represented in the computer-generated experience, the viewpoint of the currently displayed view of the three-dimensional environment changes according to the movement of the user and / or display-generated elements in the physical environment. The user perceives the sound as coming from the virtual position of the virtual sound source and experiences the visual content of the three-dimensional environment in the same reference frame based on the physical environment represented in the computer-generated experience.
[0116] In some embodiments, the reference frame is a reference frame based on a virtual three-dimensional environment of a computer-generated experience provided through the display generation components of a computer system. In some embodiments, when the reference frame is based on a virtual three-dimensional environment (e.g., an environment such as a virtual three-dimensional movie, a three-dimensional game, or a virtual office), one or more perceived sound sources have their respective spatial positions within the virtual three-dimensional environment. In some embodiments, as one or more audio output devices move around the physical environment, the audio output from the one or more audio output devices is adjusted to continue emitting sound as if the sound were coming from one or more perceived sound sources at their respective spatial positions in the virtual three-dimensional environment. If one or more perceived sound sources are moving sound sources that move through a set of spatial positions around the virtual three-dimensional environment, the audio output from the one or more audio output devices is adjusted to continue emitting sound as if the sound were coming from one or more perceived sound sources at their respective spatial positions in the virtual three-dimensional environment. In some embodiments, when audio content is output using a spatial audio output mode and reference frame based on a three-dimensional environment of a computer-generated experience, the viewpoint of the currently displayed view of the three-dimensional environment changes according to the movement of the user and / or display-generated elements in the physical environment. The user perceives the sound as coming from the virtual position of the virtual sound source and experiences the visual content of the three-dimensional environment within the same reference frame. In some embodiments, when audio content is output using a spatial audio output mode and reference frame based on a three-dimensional environment of a computer-generated experience, the viewpoint of the currently displayed view of the three-dimensional environment changes according to a movement request provided by the user and / or according to the movement of the user and / or display-generated elements in the physical environment. The user perceives the sound as coming from the virtual position of the virtual sound source and experiences the visual content of the three-dimensional environment within the same reference frame, with the user's virtual position linked to the viewpoint of the currently displayed view.
[0117] In some embodiments, the reference frame for the spatial audio output mode is fixed to an electronic device, such as a display generation component that outputs visual content corresponding to audio content being output via an audio output device (e.g., sound follows the display generation component). For example, the location of the simulated source of sound in the physical environment moves in accordance with the movement of the display generation component in the physical environment, but not with the movement of the audio output device in the physical environment. For example, in some embodiments, the display generation component is a head-mounted display or a handheld display device, and the audio output device is located in the physical environment and does not follow the user's movement. In some embodiments, the reference frame for the spatial audio effect is fixed to the display generation component and indirectly fixed to the user as the display generation component and the user move around the physical environment relative to the audio output device(s). In some embodiments, when audio content is output using the spatial audio output mode and reference frame based on the three-dimensional environment of the computer-generated experience, the viewpoint of the currently displayed view of the three-dimensional environment changes according to the movement request provided by the user and / or according to the movement of the user and / or display generation component in the physical environment. The user perceives sound as originating from a virtual position of a virtual sound source, and experiences visual content in a three-dimensional environment within the same reference frame, linking the user's virtual position to the viewpoint of the currently displayed view.
[0118] In some embodiments, at least some reference frames of spatial sound effects are fixed to the viewpoint of the currently displayed view of the three-dimensional environment (e.g., an augmented reality environment, a mixed reality environment, a virtual reality environment, etc.) presented via the display generation component. In some embodiments, the viewpoint moves relative to the three-dimensional environment to provide views of the three-dimensional environment from different positions or viewpoints within the three-dimensional environment during the computer-generated experience. In some embodiments, the viewpoint remains stationary in the three-dimensional environment during the computer-generated experience. In some embodiments, the movement of the viewpoint in the three-dimensional environment is caused by and corresponds to the movement of the display generation component in the physical environment. In some embodiments, the movement of the viewpoint in the three-dimensional environment is caused by and corresponds to the movement of the entire user relative to the physical environment or the movement of the user from head to torso. In some embodiments, the movement of the viewpoint in the three-dimensional environment is caused by and corresponds to navigation or movement requests provided by the user and / or generated by the computer system. In some embodiments, one or more perceived sound sources have their respective spatial locations within the three-dimensional environment relative to the viewpoint. For example, in some embodiments, the computer-generated experience includes visual counterparts of perceived sound sources having respective positions in a three-dimensional environment relative to the viewpoint (e.g., virtual objects, virtual light, virtual surfaces, etc., that generate sound in the computer-generated experience). In some embodiments, the computer-generated experience includes sound that does not include visual counterparts (e.g., remote or hidden virtual objects, virtual wind, sound effects, external narrator, etc., that generate sound in the computer-generated experience), but has a source corresponding to each position in the three-dimensional environment relative to the viewpoint. In some embodiments, as the viewpoint moves around the three-dimensional environment, the audio output from the audio output device(s) is adjusted to continue emitting sound as if the sound were coming from one or more perceived sound sources at each position in the three-dimensional environment.
[0119] In some embodiments, the computing system is configured to display the visual components of CGR content via a display generation component having two or more immersion levels. In some embodiments, the computer system displays the visual components of CGR content having at least a first, second, and third immersion level. In some embodiments, the computer system displays the visual components of CGR content having at least two immersion levels, providing a less immersive visual experience and a more immersive visual experience relative to each other. In some embodiments, the computing system transitions the visual content displayed via the display generation component between different immersion levels in response to a series of one or more events (e.g., the natural progression of an application or experience, the start, termination, and / or pause of an experience in response to user input, a change in the immersion level of the experience in response to user input, a change in the state of the computing device, a change in the external environment, etc.). In some embodiments, the first, second, and third immersion levels correspond to an increase in the amount of virtual content present in the CGR environment and / or a decrease in the amount of representation of the surrounding physical environment present in the CGR environment (e.g., a representation of the portion of the physical environment in front of the first display generation component). In some embodiments, the first, second, and third immersion levels correspond to different modes of content display having increased image fidelity (e.g., increased pixel resolution, increased color resolution, increased saturation, increased brightness, increased opacity, increased image detail, etc.) and / or spatial range (e.g., angular range, spatial depth, etc.) for the visual components of computer-generated content, and / or decreased image fidelity and / or spatial range for the representation of the surrounding physical environment. In some embodiments, the first immersion level is a pass-through mode in which the physical environment is fully visible to the user through the display-generated elements (e.g., as a camera view of the physical environment, or through transparent or translucent portions of the display-generated elements).In some embodiments, the visual CGR content presented in pass-through mode includes a pass-through view of the physical environment having a minimum number of virtual elements that are simultaneously visible as a view of the physical environment, or having only virtual elements around the user's view of the physical environment (e.g., indicators and controls displayed in the peripheral area of the display). For example, the view of the physical environment occupies the central and majority area of the field of view provided by the display-generating component, and only a few controls (e.g., movie title, progress bar, playback controls (e.g., play button)) are displayed in the peripheral area of the field of view provided by the display-generating component. In some embodiments, the first level of immersion is a pass-through mode in which the physical environment is fully visible to the first user through the display-generating component (e.g., as a camera view of the physical environment, or through the transparent portion of the display-generating component), and the visual CGR content is displayed in a virtual window or virtual frame that overlays, replaces, or blocks a portion of the representation of the physical environment. In some embodiments, the second level of immersion is a mixed reality mode in which a pass-through view of the physical environment is augmented with computer-generated virtual elements that occupy the central and / or majority of the user's field of view (e.g., virtual content is integrated with the physical environment within the view of the computer-generated environment). In some embodiments, the second level of immersion is a mixed reality mode in which a pass-through view of the physical environment is augmented with virtual windows, viewports, or frames that overlay, replace, or block a portion of the representation of the physical environment, and have additional depth or spatial range revealed when the display-generated elements are moved relative to the physical environment. In some embodiments, the third level of immersion is an augmented reality mode in which virtual content is displayed in a three-dimensional environment along with a representation of the physical environment, and virtual objects are distributed throughout the three-dimensional environment in positions corresponding to different locations in the physical environment. In some embodiments, the third level of immersion is a virtual reality mode in which virtual content is displayed in a three-dimensional environment without a representation of the physical environment.In some embodiments, the different immersion levels described above represent an increase in immersion levels relative to one another.
[0120] As described herein, according to some embodiments, a computer system selects an audio output mode for outputting audio content of a computer-generated experience (e.g., an application, movie, video, game, etc.) according to the level of immersion at which the visual content of the computer-generated experience is displayed by the display-generating components. In some embodiments, as the level of immersion at which the visual content is displayed increases (e.g., from a first level of immersion to a second level of immersion, from a first level of immersion to a third level of immersion, or from a second level of immersion to a third level of immersion, etc.), the computer system switches the audio output mode from a lower level of immersion output mode to a higher level of immersion output mode (e.g., from a first audio output mode to a second audio output mode, or from a first audio output mode to a third audio output mode, or from a second audio output mode to a third audio output mode, etc., where the first, second, and third audio output modes correspond to audio outputs with increasing levels of immersion). As described herein, spatial audio output modes correspond to higher levels of immersion than stereo audio output modes and mono audio output modes. Spatial audio output modes correspond to higher levels of immersion than surround sound output modes. Surround audio output modes are modes with higher levels of immersion than stereo audio output modes and mono audio output modes. Stereo audio output modes correspond to higher levels of immersion than mono audio output modes. In some embodiments, the computer system selects an audio output mode from several available audio output modes, such as mono audio output mode, stereo audio output mode, surround sound output mode, spatial audio output mode, etc., based on the level of immersion to which the visual content of the computer-generated experience is provided via the display generation component.
[0121] Figures 7A and 7B illustrate exemplary scenarios in which a first computer-generated experience is provided by a computer system (e.g., computing system 101 in Figure 1 or computing system 140 in Figure 4) that communicates with a display generation component (e.g., display 7100, another type of display generation component such as a head-mounted display) and one or more audio output devices.
[0122] In Figure 7A, the visual content of the computer-generated experience (e.g., content 7106, or other content) is presented at a first immersion level, which is a lower immersion level of two or more immersion levels capable of providing the computer-generated experience. In Figure 7B, the visual content of the computer-generated experience (e.g., content 7106, or other content) is presented at a second immersion level, which is a higher immersion level of two or more immersion levels capable of providing the computer-generated experience.
[0123] In some embodiments, one of the scenarios shown in Figures 7A and 7B may occur at the point when the computer-generated experience is initiated (e.g., in response to a user command, in response to an event generated by the computer system, etc.) without requiring a transition from a scenario shown in the other figure (e.g., without requiring the initial display of visual content with a different level of immersion). As a result, the computer system selects a corresponding audio output mode and outputs the audio content of the computer-generated experience according to the level of immersion to which the visual content of the computer-generated experience is provided.
[0124] In some embodiments, the computer system transitions from the scenario shown in Figure 7A to the scenario shown in Figure 7B, or vice versa (for example, in response to user commands, in response to events generated by the computer system, or according to pre-set conditions that are met). As a result, the computer system transitions from one audio output mode to another in response to changes in the level of immersion in which the visual content of the computer-generated experience is provided.
[0125] In some embodiments, a computer-generated experience (e.g., a 3D movie, a virtual reality game, a video, a 3D environment including user interface objects, etc.) is a virtual experience that takes place in a virtual 3D environment. In some embodiments, a computer-generated experience is an augmented reality experience that includes representations of a physical environment and virtual content. In Figures 7A and 7B, according to some embodiments, objects (e.g., object 7104, etc.) and surfaces (e.g., vertical surfaces 7004' and 7006', horizontal surface 7008', etc.) may represent virtual objects and surfaces in a virtual 3D environment (e.g., environment 7102, or another virtual environment, etc.). In Figures 7A and 7B, the 3D environment 7102 may also represent an augmented reality environment that includes representations of virtual objects and surfaces (e.g., object 7104, the surface of a virtual table, etc.) and physical objects and surfaces (e.g., vertical walls represented by representations 7004' and 7006', floors, tables, windows represented by representation 7008', etc.) according to some embodiments. In this example, environment 7102 is an environment that can exist independently of and before the display of the computer-generated experience visual content 7106.
[0126] As shown in Figure 7A, the spatial relationship between the display generation component (e.g., display 7100, or another type of display) and the user is such that the user is in a position to view the visual CGR content presented through the display generation component. For example, the user is facing the display side of the display generation component. In some embodiments, the display generation component is the display of the HMD, and the spatial relationship shown in Figure 7A corresponds to a user wearing or holding the HMD with the display side of the HMD facing the user's eyes. In some embodiments, the user is in a position to view the CGR content presented through the display generation component when the user is facing a portion of the physical environment illuminated by the projection system of the display generation component. For example, virtual content is projected onto a portion of the physical environment, and the virtual content and the portion of the physical environment are seen by the user through a camera view of the portion of the physical environment or through a transparent portion of the display generation component when the user is facing the display surface of the display generation component. In some embodiments, the display generation component emits light that forms an image on the user's retina when the user is facing the display surface of the display generation component. For example, virtual content is displayed by an LCD or LED display, superimposed on or replacing a portion of the view of the physical environment displayed by the LCD or LED display, so that a user facing the display surface of the LCD or LED display can see the virtual content together with the view of the portion of the physical environment. In some embodiments, the display generating component displays a camera view of the physical environment in front of the user, or includes a transparent or translucent portion of the portion of the physical environment in front of a first user that is visible to the user.
[0127] In some embodiments, the computer system controls one or more audio output devices that provide the user with audio output (e.g., the audio portion of the CGR content accompanying the visual portion of the displayed CGR content, system-level sound outside the CGR content, etc.). In some embodiments, the computer system generates and / or adjusts the audio output before outputting the audio CGR content using the individual audio output modes of the audio output devices, including two or more stereo audio output modes, surround sound output modes, and spatial audio output modes, corresponding to different immersion levels to which the audio CGR content can be output. In some embodiments, the computing system optionally partially or completely shields the user from sound propagating from the surrounding physical environment (e.g., via one or more active or passive noise suppression or cancellation components). In some embodiments, the amount of active sound shielding or sound passthrough is determined by the computing system based on the current immersion level related to the CGR content shown via the display generation component (e.g., no shielding in passthrough mode, partial shielding in mixed reality mode, and complete shielding in virtual reality mode, etc.).
[0128] In some embodiments, as shown in Figure 7A, the computing system displays visual CGR content 7106 via a display generation component 7100 (e.g., in response to a user command 7112 to display CGR content in a frame or viewport (e.g., a frame or viewpoint 7110, a window, a virtual screen, etc.), or in response to a transition from a lower immersion mode or a transition from a higher immersion mode (e.g., as shown in Figure 7B), etc.). At the moment shown in Figure 7A, the computing system is displaying a movie (e.g., a three-dimensional movie, a two-dimensional movie, an interactive computer-generated experience, etc.). The movie is displayed in a frame or viewpoint 7110 so that the movie content is visible simultaneously with a representation of the physical environment within the environment 7102. In some embodiments, this display mode corresponds to a low or intermediate level of immersion associated with the CGR content presented via the display generation component.
[0129] In some embodiments, the representation of a physical environment shown in a three-dimensional environment (e.g., environment 7102, another environment, etc.) includes a camera view of the portion of the physical environment that is within the first user's field of view if the user's eyes are not blocked by the presence of a display generation component (e.g., if the first user is not wearing an HMD or is holding an HMD in front of their eyes). In the display mode shown in Figure 7A, the CGR content 7106 (e.g., a movie, a three-dimensional augmented reality environment, a user interface, virtual objects, etc.) is displayed so as to overlay or replace part, but not all, of the representation of the physical environment. In some embodiments, the display generation component includes a transparent portion that allows the first user to see a portion of the physical environment. In some embodiments, in the display mode shown in Figure 7A, the CGR content 7106 (e.g., a movie, a three-dimensional augmented reality environment, a user interface, virtual objects, etc.) is projected onto a physical surface or an empty space within the physical environment and can be seen through the transparent portion of the display generation component having the physical environment, or through a camera view of the physical environment provided by the first display generation component. In some embodiments, the CGR content 7106 is displayed so as to overlay a limited portion of the display, blocking the view of a limited but not all portion of the physical environment visible through the transparent or semi-transparent portion of the first display-generating component. In some embodiments, as shown in Figure 7A, the visual CGR content is limited to a lower portion of the field of view provided by the display-generating component, such as a virtual window 7110, a virtual viewport, a virtual screen, or a position corresponding to a location on a finite physical surface, and the field of view simultaneously includes other lower portions of the three-dimensional environment, such as representations of virtual objects and / or the physical environment.
[0130] In some embodiments, as shown in Figure 7A, other user interface objects related to and / or unrelated to the CGR content (e.g., the playback control 7108, the dock with application icons, etc.) are optionally displayed simultaneously with the visual CGR content in the three-dimensional environment. In some embodiments, the visual CGR content is optionally three-dimensional content, and the viewpoint of the currently displayed view of the three-dimensional content in window 7110 moves in response to display generation components in the physical environment or user input and / or movement of the user's head.
[0131] In some embodiments, the location of the lower part of the three-dimensional environment to which the visual CGR content is limited (e.g., window 7110, viewport, etc.) is movable while the visual CGR content is being displayed. For example, according to some embodiments, the window 7110 or viewport displaying the visual CGR content is movable according to the user's pinch-and-drag gestures. In some embodiments, the window or viewport displaying the visual CGR content remains in a predetermined portion of the field of view provided by the display generating component (e.g., the center of the field of view, or a position selected by the user) when the user moves the display generating component relative to the physical environment (e.g., when the user is wearing an HMD and walking in the physical environment, or when the user is moving a handheld display in the physical environment).
[0132] In this example, as shown in Figure 7A, when displaying visual CGR content with a low or intermediate level of immersion, the computer system selects an audio output mode corresponding to the low or intermediate level of immersion, such as a stereo audio output mode, which is the output sound relative to a reference frame associated with the location of one or more audio output devices in the physical environment. In this example, according to some embodiments, the audio output devices are optionally movable relative to the display generation components and / or user in the physical environment. According to some embodiments, the audio CGR content output according to the stereo audio output mode does not take into account the position and / or movement of the window 7110 or viewport of the visual CGR content in the three-dimensional environment 7106. According to some embodiments, the audio CGR content output according to the stereo audio output mode does not take into account the position and / or movement of one or more virtual sound sources in the window 7110 or viewport of the visual CGR content. According to some embodiments, the audio CGR content output according to the stereo audio output mode does not take into account the position and / or movement of the viewpoint of the visual CGR content in the three-dimensional environment 7106. According to some embodiments, the audio CGR content output according to the stereo audio output mode does not take into account the position and / or movement of the display generation components in the physical environment. According to some embodiments, the audio CGR content output according to the stereo audio output mode is optionally locked to a reference frame associated with the user's head location when the user moves relative to the display generation components, when the user's virtual position moves relative to the three-dimensional environment represented by the CGR content (e.g., causing a movement of the viewpoint), when the window 7110 moves within the three-dimensional environment, and / or when the visual embodiment of the virtual sound source moves within the window 7110.
[0133] In some embodiments, as shown in Figure 7A, low or intermediate levels of immersion also correspond to partial shielding or partial pass-through of sound propagating from the physical environment (e.g., a portion of the physical environment surrounding the first user).
[0134] Figure 7B shows the same portion of the visual CGR content 7106 displayed by a display generation component (e.g., display 7100, or another type of display such as an HMD) using a higher level of immersion than that shown in Figure 7A. In some embodiments, switching between immersion levels can be done at any point selected by the user or computer system during the presentation of the visual CGR content. At this point, the CGR content 7106 is still displayed in the augmented reality environment 7102, but occupies a larger spatial range than that shown in Figure 7A. For example, virtual objects 7106-1, 7106-2, 7106-3, and 7106-4 within the visual CGR content 7106 are displayed in spatial positions corresponding to their physical locations within the physical environment and are integrated into the representation of the physical environment. In some embodiments, additional virtual objects, such as virtual shadows 7106-1', 7106-4', 7106-3', etc., support the virtual objects 7106-1, 7106-4, and 7106-3, etc., in the three-dimensional environment, or are added to their respective virtual positions corresponding to physical locations (e.g., physical surface locations) below the virtual objects 7106-1, 7106-4, and 7106-3, etc. In some embodiments, in accordance with the movement of the display generation components relative to the physical environment, the computing system updates the view of the three-dimensional environment 7102 in the visual CGR content 7106 in Figure 7B, as well as the viewing angle and viewing distance of the virtual objects.
[0135] In some embodiments, Figure 7B optionally represents the display of CGR content 7106 with a higher level of immersion, such as a virtual reality mode that does not have a representation of a physical environment (e.g., a 3D movie or game environment). In some embodiments, the switching performed by the computing system responds to a request from a first user (e.g., a gesture input that meets a pre-defined criterion for changing the level of immersion of the CGR content, or an event generated by the computer system based on the current context).
[0136] In some embodiments, as shown in Figure 7B, the computing system displays the visual CGR content 7106 via the display generation component 7100 (for example, in response to a user command 7114 to display the CGR content 7106 in augmented reality mode across the entire representation of the physical environment, or in response to a transition from a lower immersion mode (e.g., as shown in Figure 7A), or a transition from a higher immersion mode (e.g., virtual reality mode)). In some embodiments, as shown in Figure 7B, when displaying the CGR content 7106 using a higher level of immersion compared to that in Figure 7A, the visual CGR content 7106 is no longer limited to a limited lower portion of the field of view provided by the display generation component, such as a virtual window 7110, a virtual viewport, a location on a finite physical surface, or a virtual screen, but is distributed to different positions across different parts of the three-dimensional environment 7102. In some embodiments, other user interface objects related to and / or unrelated to the CGR content (e.g., playback control 7108, dock with application icons, etc.) are optionally displayed simultaneously with the visual CGR content 7106 in the three-dimensional environment 7102 (e.g., peripheral portion of the field of view, portion selected by the user, etc.). In some embodiments, if the visual CGR content 7106 is three-dimensional content, the viewpoint of the currently displayed view of the three-dimensional content is optionally moved in response to display generation components in the physical environment or user input and / or movement of the user's head.
[0137] In this example, as shown in Figure 7B, when displaying visual CGR content 7106 with an increased level of immersion, the computer system selects an audio output mode that corresponds to the increased level of immersion, such as a surround sound audio output mode or a spatial audio output mode that is output against a reference frame that is no longer tied to the location of one or more audio output devices in the physical environment.
[0138] In this example, according to some embodiments, the audio output device is optionally movable relative to the display generation components and / or the user in the physical environment. According to some embodiments, the audio CGR content output according to the spatial audio output mode takes into account the position and / or movement of the virtual sound source in the three-dimensional environment 7102. According to some embodiments, the audio CGR content output according to the spatial audio output mode takes into account the position and / or movement of the viewpoint of the visual CGR content in the three-dimensional environment 7106. According to some embodiments, the audio CGR content output according to the spatial audio output mode takes into account the position and / or movement of the display generation components in the physical environment.
[0139] In some embodiments, higher levels of immersion also correspond to increased sound shielding or reduced sound pass-through from the physical environment (e.g., parts of the physical environment surrounding the first user).
[0140] In some embodiments, to achieve the necessary adjustments for outputting audio CGR content in a spatial audio output mode that takes into account the movement in each environment, such as display generation components, users, audio output devices, viewpoints, and / or virtual sound sources, the computer system optionally utilizes one or more additional audio output components to output audio, compared to those used in stereo audio output mode, while continuing to reflect the position(s) and / or movement of the sound sources(s) in each reference frame(s) separated from the location of the audio output device(s). In some embodiments, the additional audio output components are located in different locations than those used in stereo audio output mode. In some embodiments, the computer system dynamically selects which audio output components to activate when outputting individual portions of audio CGR content in spatial audio output mode, based on the position and movement of the virtual sound sources in the corresponding portions of the visual CGR content of the computer-generated experience, which are simultaneously provided through display generation components with a higher level of immersion. In some embodiments, the audio output components used to output audio CGR content in spatial audio output mode are a superset of the audio output components used to output audio CGR content in stereo audio output mode and / or surround sound output mode. In some embodiments, the audio output component used to output audio CGR content in spatial audio output mode extends to a wider spatial area than the audio output component used to output audio CGR content in stereo audio output mode and / or surround sound audio output mode.
[0141] In some embodiments, the spatial audio output mode provides sound localization based on visual content, while the stereo audio output provides headlock sounds. In some embodiments, the display generation component and the audio output device are enclosed within the same head-mounted device. In some embodiments, the display generation component and the audio output device are positioned separately relative to the user's head (e.g., eyes and ears in a physical environment away from the user). In some embodiments, the display generation component is not fixedly positioned relative to the user's head, while the audio output device(s) are fixedly positioned relative to the user's ears during the presentation of CGR content. In some embodiments, the display generation component is fixedly positioned relative to the user's head during the presentation of CGR content, while the audio output device(s) are not fixedly positioned relative to the user. In some embodiments, the computer system adjusts the generation of sound corresponding to the audio CGR content while the audio CGR content is being output using a spatial audio output mode, in accordance with the relative movement and spatial configuration of the display generation components, the user, and the audio output device(s), to provide sound localization based on the visual content (e.g., movement of viewpoint, change of virtual sound source, movement of virtual sound source, etc.).
[0142] In some embodiments, when providing sound localization based on the position of virtual sound sources within visual CGR content, the computer system determines the virtual position of individual virtual sound sources within the three-dimensional environment of the CGR content, determines an appropriate reference frame for the sound corresponding to the individual virtual sound source (e.g., a reference frame based on the physical environment, virtual environment, viewpoint, etc., selected based on the type of CGR content being presented), determines the individual position of the virtual sound source within the selected reference frame based on the current position of the individual sound source within the three-dimensional environment of the CGR content, and controls the operation of the audio output component of the audio output device(s) to output the sound corresponding to the individual sound source(s) so that the sound is perceived as originating from the individual position of the individual sound source(s) within the selected reference frame(s) in the physical environment. In the example shown in Figure 7B, if the virtual object 7106-1 is a virtual sound source (e.g., a virtual bird, a virtual train, a virtual assistant, etc.) associated with an audio output (e.g., a chirp, a training chugging sound, a speech sound, etc.), and the audio CGR content is output using spatial audio output mode, the computer system optionally controls the audio output component of the virtual sound source's output so that, when perceived by the user, the sound appears to originate from a physical location corresponding to the current virtual position of the virtual object 7106-1 in the three-dimensional environment 7102, regardless of the movement of the display generation components, the user's movement, and / or the movement of the audio output device(s) in the physical environment.Similarly, in the example shown in Figure 7B, if the virtual object 7106-3 is another virtual sound source (e.g., another virtual bird, a virtual conductor, etc.) associated with another audio output (e.g., another chirp, a whistle, etc.), when the audio CGR content is output using spatial audio output mode, the computer system optionally controls the audio output component of the output of this other virtual sound source so that, when perceived by the user, the sound appears to originate from a physical location corresponding to the current virtual position of the virtual object 7106-3 in the three-dimensional environment 7102, regardless of the movement of the display generation components, the user's movement, and / or the movement of the audio output device(s) in the physical environment.
[0143] In some embodiments, when providing sound localization based on the user's position, the computer system determines the virtual position of individual virtual sound sources in the three-dimensional environment of the CGR content, determines a reference frame associated with the user's location relative to the three-dimensional environment of the CGR content, determines the individual position of the virtual sound sources in the reference frame based on the user's location, and controls the operation of the audio output component of the audio output device(s) to output sound corresponding to the individual sound sources so that the sound is perceived as originating from the individual position of the individual sound source in a reference frame fixed to the user's current location in the physical environment. In the example shown in Figure 7B, the virtual sound sources associated with the audio output (e.g., external narrator, virtual assistant, ambient sound source, etc.) optionally do not have corresponding virtual objects. When the audio CGR content is output using spatial audio output mode, the computer system optionally controls the audio output component of the virtual sound source's audio output so that, when perceived by the user, the sound appears to originate from a fixed location or region relative to the user, regardless of the movement of the display generation components, the user's movement, and / or the movement of the audio output device(s) in the physical environment. According to some embodiments, the viewpoint of the visual CGR content changes optionally according to the movement of the display generation components and / or the user's movement, while the audio output corresponding to the virtual sound source remains fixed for the user.
[0144] Figures 7C to 7H are block diagrams illustrating, in several embodiments, how the appearance of parts of virtual content may change when important physical objects approach the display generation components or the user's location (for example, allowing a representation of part of a physical object to break through the virtual content, or changing one or more visual properties of the virtual content based on the visual properties of part of a physical object).
[0145] In some embodiments, when displaying virtual content in a three-dimensional environment (e.g., environment 7126 in Figures 7C-7H, another environment, etc.) (e.g., a virtual reality environment, an augmented reality environment, etc.), all or part of the view of the physical environment is blocked or replaced by the virtual content (e.g., virtual objects 7128, 7130 in Figure 7D, etc.). In some cases, it is advantageous to give display priority to certain physical objects in the physical environment (e.g., scene 105 in Figures 7C, 7E, and 7G) over virtual content, such that at least some of the physical objects (e.g., physical object 7122 in Figure 7C, another physical object important to the user, etc.) are visually represented in the view of the three-dimensional environment (e.g., as shown in Figures 7F and 7H). In some embodiments, the computer system utilizes various criteria for determining whether to give display priority to an individual physical object so that the representation of the individual physical object can overwrite the portion of the virtual content currently displayed in the three-dimensional environment, when the location of the individual physical object in the physical environment corresponds to the position of a portion of the virtual content in the three-dimensional environment. In some embodiments, the criteria include the requirement that at least a portion of the physical object has entered an approaching threshold spatial region (e.g., spatial region 7124 in Figures 7C, 7E, and 7G, another spatial region, etc.) surrounding the user of the display-generating component (e.g., user 7002 viewing the virtual content through the display-generating component, user whose view of the portion of the physical object is blocked or replaced by the display of the virtual content, etc.), and the additional requirement that the computer system detects the presence of one or more characteristics of the physical object (e.g., physical object 7122 in Figure 7C, another physical object important to the user, etc.) that indicate the increased importance of the physical object to the user.In some embodiments, a physical object of increased importance to the user may be (for example, as shown in the examples in Figures 7C to 7H) the user's friends or family, the user's team members or supervisors, the user's pet, etc. In some embodiments, a physical object of increased importance to the user may be a person or object that requires the user's attention to deal with an emergency. In some embodiments, a physical object of greater importance to the user may be a person or object that requires the user's attention to take action that the user does not miss. The criteria can be adjusted by the user based on the user's needs and desires, and / or by the system based on contextual information (e.g., time, location, scheduled events, etc.). In some embodiments, giving display priority to physical objects that are more important than virtual content and visually representing at least a portion of the physical object in a view of the three-dimensional environment includes replacing the display of a portion of the virtual content (e.g., a portion of virtual object 7130 in Figure 7F, a portion of virtual object 7128 in Figure 7H, etc.) with a representation of a portion of the physical object, or changing the appearance of a portion of the virtual content in accordance with the appearance of a portion of the physical object. In some embodiments, even if the position corresponding to the location of a part of a physical object is visible within the field of view provided by a display generation component (e.g., the position is currently occupied by virtual content), at least a portion of the physical object (e.g., the ears and body of pet 7122 in Figure 7F, or a portion of pet 7122's body in Figure 7H) is not visually represented in the view of the three-dimensional environment and remains blocked or replaced by the display of virtual content.In some embodiments, portions of the three-dimensional environment that have been modified to indicate the presence of physical objects, and portions of the three-dimensional environment that have not been modified to indicate the presence of physical objects (for example, portions of the three-dimensional environment (e.g., parts of virtual object 7128, virtual object 7130, etc.) may continue to change based on the progress of the computer-generated experience and / or user interaction with the three-dimensional environment, etc.) correspond to positions on virtual objects or continuum portions of surfaces (e.g., parts of virtual object 7128, virtual object 7130, etc.).
[0146] In some embodiments, when a user is engaged in a computer-generated experience, such as a virtual reality or augmented reality experience, through a display-generated component, the user's view of the physical environment is blocked or obscured by the presence of virtual content in the computer-generated experience. In some embodiments, while a user is engaged in a virtual reality or augmented reality experience, it is desirable to make known or visually indicate to the user the presence of important physical objects (e.g., people, pets, etc.) that are close to the user's physical vicinity. In some embodiments, important physical objects are within the user's potential field of view, but with respect to the presence of the display-generated component and the virtual content of the computer-generated experience (e.g., if the display-generated component and / or virtual content are not present, the physical object is visible to the user), a portion of the virtual content positioned to correspond to a first portion of the physical object is removed or altered to reflect the appearance of the first portion of the physical object, while another portion of the virtual content positioned to correspond to another portion of the physical object adjacent to the first portion of the physical object is not removed or altered to reflect the appearance of the other portion of the physical object. In other words, the virtual content is removed or altered gradually, part by part, to facilitate the disruption of the computer-generated experience, rather than being removed or altered abruptly to show all portions of a physical object that are potentially within the user's field of view.
[0147] In various embodiments, important physical objects are identified by a computer system based on criteria including at least one requirement independent of or unrelated to the distance between the physical object and the user. In some embodiments, when the computer system determines whether an approaching physical object is an important physical object for the user, it considers various pieces of information such as the user's previously entered settings, the presence of previously identified characteristics, the current context, and the presence of marker objects or signals associated with the physical object, to ensure that the computer-generated experience is not visually disruptive.
[0148] As shown in Figure 7C, user 7002 is present in a physical environment (e.g., scene 105, or another physical environment). User 7002 is in a predetermined position relative to a display generation component (e.g., display generation component 7100, another type of display generation component such as an HMD) in order to view content displayed via the display generation component. The predefined spatial region 7124 surrounding user 7002 is shown in Figure 7C by a dashed line surrounding user 7002. In some embodiments, the predefined spatial region 7124 is a three-dimensional region surrounding user 7002. In some embodiments, the predefined spatial region 7124 is defined by a predefined threshold distance (e.g., arm length, 2 meters) to characteristic locations of the user in the physical environment (e.g., location of the user's head, location of the user's center of gravity). In some embodiments, the predefined spatial region 7124 has a boundary surface that is greater from the front of the user (e.g., face, chest, etc.) than from the back of the user (e.g., back of the head, etc.). In some embodiments, the predefined spatial region 7124 has a boundary surface that is at a greater distance from one side of the user than from the other side of the user (for example, greater distance from the left side of the user than from the right side of the user, or vice versa). In some embodiments, the predefined spatial region 7124 has a boundary surface that is symmetrical on both sides of the user. In some embodiments, the predefined spatial region 7124 is at a greater distance from the upper part of the user's body (e.g., the user's head, the user's chest, etc.) than from the lower part of the user's body (e.g., the user's feet, the user's legs, etc.). In some embodiments, the display generation component has a fixed spatial relationship with the user's head. In some embodiments, the display generation component surrounds the user's eyes and blocks the user's view of the physical environment except for the view provided through the display generation component.
[0149] In some embodiments, as shown in Figure 7C, the physical environment contains other physical objects (e.g., physical object 7120, physical object 7122, etc.) and physical surfaces (e.g., walls 7004, 7006, floor 7008, etc.). In some embodiments, at least some of the physical objects are stationary objects relative to the physical environment. In some embodiments, at least some of the physical objects move relative to the physical environment and / or the user. In the example shown in Figure 7C, physical object 7122 represents an instance of a first type of physical object that is important to user 7002 based on an evaluation according to a predetermined criterion. Physical object 7120 represents an instance of a second type of physical object that is not important to user 7002 based on an evaluation according to a predetermined criterion. In some embodiments, the physical environment may contain only one of the two types of physical objects at a given time. In some embodiments, after the user 7002 has already started the computer-generated experience, one of the two types of physical objects may enter the physical environment, and the user may not necessarily perceive the physical object entering the physical environment due to the presence of display generation components and / or virtual content displayed through the display generation components.
[0150] Figure 7D shows that the display generation component displays a view of the three-dimensional environment 7126 at a time corresponding to the view shown in Figure 7C. In this example, the three-dimensional environment 7126 is a virtual three-dimensional environment that does not include representations of the physical environment surrounding the display generation component and the user. In some embodiments, the virtual three-dimensional environment includes virtual objects (e.g., virtual object 7128, virtual object 7130, user interface objects, icons, avatars, etc.) and virtual surfaces (e.g., virtual surfaces 7132, 7136, and 7138, virtual windows, virtual screens, user interface background surfaces, etc.) at various positions within the virtual three-dimensional environment 7126. In some embodiments, the movement of the user and / or the display generation component changes the viewpoint of the currently displayed view of the three-dimensional environment 7126 in response to the movement of the user and / or the display generation component in the physical environment. In some embodiments, the computer system moves or changes the viewpoint of the currently displayed view of the three-dimensional environment 7126 according to events generated by the computer system based on user input, the pre-programmed progression of the computer-generated experience, and / or the fulfillment of pre-set conditions. In some embodiments, virtual content (e.g., movies, games, etc.) changes over time as the computer-generated experience progresses, without user input.
[0151] In some embodiments, the three-dimensional environment 7126 shown in Figure 7D represents an augmented reality environment, where virtual content (e.g., virtual surfaces and virtual objects) is displayed simultaneously with a representation of the physical environment (e.g., Scene 105, or another physical environment surrounding the user). At least a portion of the representation of the physical environment in front of the user (e.g., one or more consecutive (or continuous) portions of the physical environment, and / or discontinuous and disconnected portions) is blocked, replaced, or hidden by the virtual content displayed by the display-generating component. For example, in some embodiments, virtual surfaces 7132 and 7136 are representations of walls 7006 and 7004 in the physical environment 105, virtual surface 7134 is a representation of floor 7008 in the physical environment 105, and virtual objects 7128 and 7130 block, replace, or overlay at least a portion of the representation of the physical environment (e.g., a portion of the representations of walls 7006 and floor 7008, and a portion of the representations of physical objects 7120 and 7122).
[0152] As shown in Figures 7C and 7D, when both physical objects 7122 and 7120 are outside the predefined spatial portion 7124 surrounding user 7002, but are within the user's potential field of view without the presence of the display generation component 7100, the virtual content of the three-dimensional environment 7126 (e.g., virtual objects 7128 and 7130, etc.) is displayed via the display generation component 7100 without the disruption of the physical objects 7122 and 7120. For example, if the three-dimensional environment 7126 is a virtual environment, even though the positions corresponding to the locations of the physical objects 7122 and 7120 are within the field of view provided by the display generation component, the portions of the virtual content having their respective virtual positions corresponding to the locations of the physical objects 7122 and 7120 are displayed normally according to the original CGR experience. In another example, if the three-dimensional environment 7126 is an augmented reality environment, portions of the virtual content having virtual positions corresponding to the locations of physical objects 7122 and 7120 will be displayed normally according to the original CGR experience, even if the positions corresponding to the locations of physical objects 7122 and 7120 are within the field of view provided by the display generation component, and even if some parts of the physical environment (e.g., parts of the walls, the floor, parts of physical objects 7122 and 7120, etc.) are visible in space that is not currently occupied or visually blocked by the virtual content of the CGR experience.
[0153] Figures 7E to 7F illustrate a scenario in which physical objects 7122 and 7120 move closer to user 7002 within the physical environment 105. In this scenario, only a portion of the total spatial range of physical object 7122 lies within a predefined spatial region 7124 surrounding user 7002. Similarly, only a portion of the total spatial range of physical object 7120 lies within a predefined spatial region 7124 surrounding user 7002. In some embodiments, in response to the detection of movement of physical objects in the physical environment (e.g., physical object 7120, physical object 7122, etc.), and based on the determination that the user is within a threshold distance of the physical object (for example, the threshold distance is determined based on the boundary of the predefined spatial region 7124, and the individual relative spatial relationships between the user and the physical object, a fixed predefined threshold distance, etc.), the computer system determines whether the physical object is important to the user according to predefined criteria.
[0154] In this example, the physical object 7122 meets the requirements for being recognized as an important physical object for user 7002, and therefore the computer system modifies the appearance of the virtual content displayed at the position corresponding to the location of the first part of the physical object 7122 according to the appearance of the first part of the physical object 7122. As shown in Figure 7F, the virtual content shown at the position corresponding to the location of the first part of the physical object 7122 is removed, and the representation 7122-1' of the first part of the physical object 7122 (e.g., part of the pet's head, the head of the physical object 7122, etc.) is revealed. In some embodiments, the visual properties (e.g., color, simulated refractive index, transparency, brightness, etc.) of the virtual content shown at the position corresponding to the location of the first part of the physical object 7122 (e.g., part of the virtual object 7130 in this example, Figure 7F) are modified according to the appearance of the first part of the physical object 7122. In some embodiments, as shown in Figure 7F, the virtual content of positions corresponding to the locations of some parts of a physical object 7122 within a preset spatial region 7124 remains unchanged in the view of the three-dimensional environment 7126 (e.g., the part of the virtual object 7130 around the wavy edge of representation 7122-1' in Figure 7F) even though those parts of the physical object (e.g., part of the head and part of the body of the physical object 7122, as shown in Figure 7E) are within the user's threshold distance and the display-generated components are removed, even though they are now within the user's natural field of view. In some embodiments, the virtual content of positions corresponding to the locations of all parts of a physical object 7122 within a preset spatial region 7124 may ultimately be removed or modified in the view of the three-dimensional environment 7126 after a period in which the parts of the physical object 7122 remain within the preset spatial region 7124.
[0155] In this example, the physical object 7120 does not meet the requirements to be recognized as an important physical object for user 7002, and therefore the computer system does not change the appearance of the virtual content (e.g., virtual object 7128 in Figure 7F) displayed at the position corresponding to the location of the first part of the physical object 7120, according to the appearance of the first part of the physical object 7120. As shown in Figure 7F, the virtual content displayed at the position corresponding to the location of the first part of the physical object 7120 is not removed, and the first part of the physical object 7120 is not visible in the view of the three-dimensional environment 7126.
[0156] In some embodiments, the contrast between the processing of physical object 7120 and physical object 7122 is based on pre-set criteria on which physical objects 7120 and 7122 are evaluated. For example, when physical object 7122 is important, physical object 7120 has not been previously marked as important by the user. When physical object 7122 is present, physical object 7120 is not moving toward the user at a threshold speed. When physical object 7122 is present, physical object 7120 is not a person or a pet. When physical object 7122 is a person but not speaking, physical object 7120 is a person but is not speaking when approaching the user. When physical object 7122 is present, physical object 7120 is not wearing a pre-set identifier object (e.g., a radio-transmitted ID, an RFID tag, a color-coded tag, etc.).
[0157] In the view shown in Figure 7F, the first portion of the physical object 7120 is within the threshold distance of the user 7002, its corresponding position in the computer-generated environment 7126 is visible to the user based on the user's field of view of the computer-generated environment, the position corresponding to the first portion of the physical object 7120 is not blocked from the user's field of view by another physical object or the position corresponding to another portion of the physical object 7120, and the computer system still does not change the appearance of the portion of virtual content displayed at the position corresponding to the first portion of the physical object 7120 (e.g., virtual object 7128 in Figure 7F) because the physical object 7120 does not meet the pre-set criteria for being an important physical object to the user 7002. For example, a ball does not meet the pre-set criteria that require the first physical object to be a person or a pet. When the ball rolls towards the user, the computer system does not change the appearance of the virtual content displayed at the position in the computer-generated environment corresponding to the ball's location relative to the user. In contrast, when the pet approaches the user, the computer system changes the appearance of the virtual content displayed at the position corresponding to the part of the pet that is within the user's preset distance, without changing the appearance of the virtual content displayed at the position corresponding to the other part of the pet that is not within the user's preset distance, even though the positions corresponding to other parts of the pet are also within the user's current field of view.
[0158] Figures 7G and 7H show that later, both physical objects 7120 and 7122 move closer to the user and fully enter the pre-defined spatial portion 7124 surrounding the user, and are within the user's field of view when the display generation components are removed.
[0159] As shown in Figure 7H, the computer system modifies the appearance of virtual content displayed at positions corresponding to the location of the second part of the physical object 7122 (e.g., the part including the first part of the physical object 7122 and additional parts of the physical object 7122 that have entered a predefined spatial area 7124 surrounding the user) in accordance with the appearance of the second part of the physical object 7122 (e.g., at least parts of virtual objects 7130 and 7128). As shown in Figure 7H, the virtual content shown at positions corresponding to the location of the second part of the physical object 7122 is removed, revealing representation 7122-2' of the second part of the physical object 7122 (e.g., a larger portion of the physical object 7122 than that corresponding to representation 7122-1' shown in Figure 7F). In some embodiments, the visual properties of the virtual content shown at positions corresponding to the location of the second part of the physical object 7122 (e.g., color, simulated refractive index, transparency, brightness, etc.) are modified in accordance with the appearance of the second part of the physical object 7122. In some embodiments, as shown in Figure 7H, the virtual content of positions corresponding to the locations of some parts of a physical object 7122 within a preset spatial region 7124 remains unchanged in the view of the three-dimensional environment 7126, even though those parts of the physical object are within the user's threshold distance and are within the user's natural field of view at this point if the display generation components are removed. In some embodiments, the virtual content of positions corresponding to the locations of all parts of a physical object 7122 within a preset spatial region 7124 may be ultimately removed or modified in the view of the three-dimensional environment 7126 after a period in which the parts of the physical object 7122 remain within the preset spatial region 7124.
[0160] In this example, the physical object 7120 does not meet the requirements to be recognized as an important physical object for user 7002, and therefore the computer system does not change the appearance of the virtual content displayed at the position corresponding to the location of the second part of the physical object 7120 according to the appearance of the second part of the physical object 7120. As shown in Figure 7H, in Figure 7H, the virtual content shown at the position corresponding to the location of the second part of the physical object 7120 is not removed, and the second first part of the physical object 7120 is not visible in the view of the three-dimensional environment 7126.
[0161] In some embodiments, there is no clear structural or visual division between the portion of the physical object 7122 revealed in the view of the three-dimensional environment 7126 and the other portion of the physical object 7122 not revealed in the view of the three-dimensional environment, which would provide a basis for different processing applied to the different portions of the first physical object. Instead, the difference is based on the fact that the revealed portion of the physical object 7120 is within the user's threshold distance or area, while the other portion of the physical object 7122 is not within the user's threshold distance or area. For example, the physical object 7122 is a pet, and at a given point in time, the portion of the physical object revealed by the removal or alteration of the appearance of the virtual content includes a first portion of the pet's head (e.g., nose, whiskers, part of the face, etc.), while the remaining portion of the physical object not revealed by the removal or alteration of the virtual content includes the rest of the pet's head (e.g., the rest of the face and ears, etc.) and additional portions of the torso connected to the head, which are not within the user's threshold distance.
[0162] In some embodiments, the portion of virtual content that is modified or deleted to reveal the presence of a portion of a physical object 7122 within a predefined spatial region 7124 is part of a contiguous virtual object or surface, while the rest of the contiguous virtual object or surface remains visible without modification. For example, as shown in Figure 7F, only a portion of the virtual object 7130 is removed or its appearance is altered in order to reveal the presence of a portion of a physical object 7122 at a location within a predefined spatial region 7124 that has a position corresponding to the position of a portion of the virtual object 7130.
[0163] In some embodiments, a physical object 7122 is detected by a computer system and qualifies as a physical object of importance to user 7002 based on first characteristics that distinguish it from a person and a non-personal physical object. In some embodiments, the first characteristics include a predefined facial structure (e.g., the presence and / or movement of eyes, the relative location of the eyes, nose, and mouth), the proportions and relative positions of body parts on the physical object 7122 (e.g., the head, body, and limbs), a human voice associated with the movement of the physical object 7122, and movement patterns associated with human walking or running (e.g., arm swing, gait). A physical object 7120 does not have the qualities of a physical object of importance to user 7002 because the first characteristics are not present in the physical object 7120.
[0164] In some embodiments, a physical object 7122 is detected by a computer system and qualifies as a physical object of importance to user 7002 based on a second characteristic indicating a human voice coming from the physical object 7122 as it moves toward the user. In some embodiments, the second characteristic includes a pre-defined vocal characteristic of the sound originating from the location of the physical object 7122 (e.g., the presence of a voiceprint, a human language utterance pattern, etc.), the characteristics of the human voice associated with the movement of the physical object 7122, and the utterance of one or more pre-defined words (e.g., "Hello!", "Hi!", "Hello!", "[Username]", etc.). A physical object 7120 does not have the quality of a physical object of importance to user 7002 because the second characteristic is not present in the physical object 7120.
[0165] In some embodiments, a physical object 7122 is detected by a computer system and qualifies as a physical object of importance to user 7002 based on a third characteristic that distinguishes the animal from humans and non-human physical objects. In some embodiments, the third characteristic includes a preset head structure (e.g., presence and / or movement of eyes, relative location of eyes, nose, ears, whiskers, and mouth), proportions and relative positions of body parts of the physical object 7122 (e.g., head, body, tail, and limbs), presence of fur, fur color and pattern, etc.), detection of animal vocalizations versus human voices in conjunction with the movement of the physical object 7122, and detection of movement patterns associated with the animal walking or running (e.g., four legs on the ground, wing flapping, walking, etc.). A physical object 7120 does not have the qualities of a physical object of importance to user 7002 because the third characteristic is not present in the physical object 7120.
[0166] In some embodiments, a physical object 7122 is detected by a computer system and qualifies as an important physical object for user 7002 based on a fourth characteristic, which is a characteristic movement speed of the physical object 7122 that exceeds a preset threshold speed. In some embodiments, the characteristic movement speed includes the movement speed of at least a portion of the physical object relative to another part of the physical object or physical environment (e.g., waving a person's hand, popping a cork to open a bottle), or the movement speed of at least a portion of the physical object toward the user. A physical object 7120 does not have the quality of an important physical object for user 7002 because its characteristic movement speed does not meet the preset threshold speed.
[0167] In some embodiments, physical object 7122 is qualified as an important physical object for user 7002 based on a fifth characteristic of physical object 7122 that is detected by a computer system and indicates the occurrence of an event requiring the user's immediate attention (e.g., an emergency, a danger, etc.). In some embodiments, the fifth characteristic includes flashing lights, motion patterns (e.g., opening or closing a door or window, a person waving their hand, etc.), vibrations (e.g., a sign shaking, curtains, falling objects, etc.), shouting, sirens, etc. Physical object 7120 does not have the quality of an important physical object for user 7002 because the fifth characteristic is not present in physical object 7120.
[0168] In some embodiments, physical object 7122 is detected by a computer system and qualifies as a physical object of importance to user 7002 based on a sixth characteristic of physical object 7122 that indicates the presence of an identifier object on the physical object (e.g., RFID, badge, ultrasonic tag, serial number, logo, name, etc.). Physical object 7120 does not have the quality of a physical object of importance to user 7002 because the sixth characteristic is not present in physical object 7120.
[0169] In some embodiments, a physical object 7122 is detected by a computer system and is deemed eligible as an important physical object for user 7002 based on a seventh characteristic of the physical object 7122, which is based on the motion patterns of the physical object (e.g., motion patterns of at least a portion of the physical object relative to another part of the physical environment or the physical object, or motion patterns of at least a portion of the physical object relative to the user). A physical object 7120 does not have the qualities of an important physical object for user 7002 because the seventh characteristic is not present in the physical object 7120.
[0170] In some embodiments, a physical object 7122 is detected by a computer system and is qualified as an important physical object for user 7002 based on an eighth characteristic of the physical object 7122, which is based on a match between recognized identification information of the physical object (e.g., spouse, favorite pet, boss, child, police, train conductor, etc.) and first pre-set identification information (e.g., identifying as previously established as "important," "requires attention," etc.) (e.g., a match or correspondence exceeding a threshold confidence value determined by a computer algorithm or artificial intelligence (e.g., facial recognition, voice recognition, etc.) based on detected sensor data, image data, etc.). A physical object 7120 does not have the quality of an important physical object for user 7002 because the eighth characteristic is not present in the physical object 7120.
[0171] Figures 7I to 7N are block diagrams illustrating the application of visual effects to regions in a three-dimensional environment corresponding to a portion of the physical environment identified based on a scan of that portion of the physical environment (characterized, for example, by shape, plane, and / or surface), according to several embodiments.
[0172] In some embodiments, the computer system displays a representation of a three-dimensional environment, including a representation of the physical environment (e.g., Scene 105 in Figure 7I, another physical environment), in response to a request to display such an environment (e.g., in response to a user putting on a head-mounted display, in response to a user request to start an augmented reality environment, in response to a user request to end a virtual reality experience, in response to a user turning on or waking up display generation components from a low-power state, etc.). In some embodiments, the computer system initiates a scan of the physical environment to identify objects and surfaces within the physical environment (e.g., walls 7004, 7006, floor 7008, object 7014, etc.) and optionally constructs a three-dimensional or pseudo-three-dimensional model of the physical environment based on the identified objects and surfaces within the physical environment. In some embodiments, the computer system initiates a scan of the physical environment in response to receiving a request to display a three-dimensional environment (e.g., if the physical environment has not been previously scanned and characterized by the computer system, or if a rescan is requested by the user or the system based on a preset rescan criterion that is met (e.g., the last scan was performed more than a threshold time ago, the physical environment has changed, etc.)). In some embodiments, the computer system initiates a scan in response to the detection of a user's hand (e.g., hand 7202 in Figure 7K) touching a part of the physical environment (e.g., a physical surface (e.g., the top surface of physical object 7014, the surface of wall 7006, etc.), a physical object, etc.). In some embodiments, the computer system initiates a scan in response to the detection of a user's line of sight (e.g., line of sight 7140 in Figure 7J, another line of sight, etc.) directed to a position corresponding to a part of the physical environment meeting a preset stability and / or duration criterion. In some embodiments, the computer system displays visual feedback (e.g., visual effects 7144 in Figures 7K to 7L) regarding the progress and results of the scan (e.g., identification of physical objects and surfaces in the physical environment, determination of the physical and spatial properties of physical objects and surfaces, etc.).In some embodiments, the visual feedback includes displaying individual visual effects (e.g., visual effect 7144) on distinct parts of the three-dimensional environment that are touched by the user's hand (e.g., the top surface of a physical object 7014) and that correspond to a portion of the physical environment identified based on a scan of that portion of the physical environment. In some embodiments, as shown in Figures 7K to 7L, the visual effects (e.g., visual effect 7144) include a representation of motion that extends from and / or propagates from a distinct part of the three-dimensional environment (e.g., a position corresponding to the touch location of the hand 7202). In some embodiments, the computer system displays the visual effects in response to the detection of the user's hand touching a distinct part of the physical environment, and the three-dimensional environment is displayed after the scan of the physical environment is complete, in response to a previous request to display the three-dimensional environment.
[0173] In some embodiments, when a computer system performs a scan of a physical environment in preparation for the generation of a mixed reality environment (e.g., an augmented reality environment, an augmented virtual environment, etc.), it may be helpful to receive user input identifying a region of interest and / or a clearly defined area of a surface or plane in order to fix the scan of the physical environment and identify objects and surfaces within the physical environment. It is also advantageous to provide the user with visual feedback regarding the progress and results of the scan and characterization of the physical environment from a position corresponding to the location of the user's input, so that if the position does not yield the correct characterization, the user can adjust the input and restart the scan from a different location or surface of the physical environment. In some embodiments, after a physical surface has been scanned and identified based on the scan, the computer system displays an animated visual effect at a position corresponding to the identified surface, and the animated visual effect starts from a position corresponding to the contact location between the physical surface and the user's hand and propagates. In some embodiments, to further confirm the location of interest, the computer system requires that gaze input be detected at the position of the physical surface that the user is touching. In some embodiments, the line of sight position does not need to overlap with the position corresponding to the user's touch location, as long as both positions are on the same extended physical surface and / or within a threshold distance of each other.
[0174] As shown in Figure 7I, user 7002 is present in a physical environment (e.g., scene 105, or another physical environment). User 7002 is in a predetermined position relative to a display generation component (e.g., display generation component 7100, another type of display generation component such as an HMD) in order to view content displayed via the display generation component. In some embodiments, the display generation component has a fixed spatial relationship with the user's head. In some embodiments, the display generation component surrounds the user's eyes and blocks the user's view of the physical environment except for the view provided through the display generation component. In some embodiments, as shown in Figure 7C, the physical environment includes physical objects (e.g., physical object 7014, and other physical objects) and physical surfaces (e.g., walls 7004 and 7006, floor 7008, etc.). The user can view different locations within the physical environment through the view of the physical environment provided through the display generation component, and the location of the user's gaze is determined by an eye-tracking device, such as the eye-tracking device disclosed in Figure 6. In this example, the physical object 7014 has one or more surfaces (for example, a horizontal top surface, a vertical surface, a plane, a curved surface, etc.).
[0175] Figure 7J displays a view 7103 of the physical environment 105 displayed via a display generation component. According to some embodiments, the view of the physical environment includes representations of physical surfaces and objects in a portion of the physical environment as seen from a viewpoint corresponding to the location of the display generation component 7100 in the physical environment (for example, if the display generation component 7100 is an HMD, the location also corresponds to the user's eyes or head). In Figure 7J, the view 7103 of the physical environment includes representations 7004' and 7006' of two adjacent walls (e.g., walls 7004 and 7006) in the physical environment for the user and the display generation component, a representation 7008' of the floor 7008, and a representation 7014' of a physical object 7014 (e.g., furniture, objects, fixtures, etc.) in the physical environment. According to some embodiments, the spatial relationship between physical surfaces and physical objects in the physical environment is represented in the three-dimensional environment by the spatial relationship between the representations of physical surfaces and physical objects in the three-dimensional environment. As the user moves the display generation component relative to the physical environment, different views of the physical environment from different viewpoints are displayed through the display generation component. In some embodiments, if the physical environment is unknown to the computer system, the computer system performs a scan of the environment to identify surfaces and planes and construct a three-dimensional model of the physical environment. After the scan, according to some embodiments, the computer system can define the position of virtual objects relative to the three-dimensional model, and as a result, the virtual objects can be placed in the mixed reality environment based on the three-dimensional model, which has various spatial relationships with the physical surfaces and object representations in the three-dimensional environment. For example, the virtual object may optionally be given an upright orientation relative to the three-dimensional model and displayed in a position and / or orientation that simulates a specific spatial relationship with the physical surface or object representation (e.g., overlapping, standing, parallel, perpendicular, etc.).
[0176] In some embodiments, as shown in Figure 7J, the computer system detects a gaze input (e.g., gaze input 7140 in this example) directed towards a portion of the representation of the physical environment within a view 7013 of the three-dimensional environment. In some embodiments, the computer system displays a visual indicator (e.g., visual indicator 7142) at the position of the gaze. In some embodiments, the position of the gaze is determined based on the user's gaze and the focal length of the user's eyes, as detected by the computer system's eye-tracking device. In some embodiments, it is difficult to highly confirm the precise location of the user's gaze before the scan of the physical environment is completed. In some embodiments, the area occupied by the representation 7014' of a physical object can be identified by two-dimensional image segmentation before the three-dimensional scan of the physical environment is performed or completed, and the location of the gaze can be determined to be the area occupied by the representation 7014', as determined by the two-dimensional segmentation.
[0177] In some embodiments, as the user moves the display-generating component around the physical environment and looks at different surfaces or objects through the display-generating component in search of a suitable position to begin scanning, the computer provides real-time feedback to show the user the location of the line of sight in the portion of the physical environment that is currently within the field of view provided by the display-generating component.
[0178] In Figures 7K to 7L, while the user's line of sight 7140 is directed towards the representation 7014' of the physical object 7014, the computer system detects that the user's hand has moved to a first location on the upper surface of the physical object 7014 within the physical environment and maintains contact with the upper surface of the physical object 7014 at the first location. In response to the detection of the user's hand 7202 in contact with the upper surface of the physical object 7014 (e.g., optionally in relation to the detection of the user's line of sight 7140 on the same surface of the physical object 7014), the computer system begins scanning the physical environment from the location of the user's hand (e.g., from the location of contact between the user's hand and the upper surface of the physical object 7014). In some embodiments, the computer system optionally performs scanning of other parts of the physical environment in addition to and in parallel with the scanning at the location of the user's hand. Once a portion of the surface of the physical object 7014 near the contact location has been scanned and characterized (e.g., as a plane or a curved surface), the computer system displays visual feedback indicating the results and progress of the scan. In Figure 7K, the appearance of a portion of representation 7014' corresponding to the location of the user's contact with the physical object 7014, and its vicinity, is altered by visual effects (e.g., highlighted, animated, and / or modified in terms of color, brightness, transparency, and / or opacity). The visual effects have one or more spatial properties (e.g., position, orientation, surface properties, spatial extent, etc.) based on the results of a scan of the user's contact location with the physical object or a portion of the physical surface near it. For example, in this case, the computer system determines that representation 7014' is a plane having a horizontal orientation at a position corresponding to the location of the tip of the user's hand 7202, based on a scan of an area near the location of the tip of the index finger (e.g., the location of contact between the user's hand 7202 and the physical object 7014). In some embodiments, the tip of the user's finger provides an anchor location for surface scanning.In some embodiments, depth data of the physical environment at the location of the user's fingertip is correlated with depth data of the user's fingertip, and the accuracy of the scan is improved by this additional constraint.
[0179] In Figures 7L to 7M, while the user's hand 7202 maintains contact with the top surface of a physical object 7014 in the physical environment, the computer system optionally continues to apply and display visual feedback 7144 at the initial touch location on the top surface of the physical object 7014 to indicate the progress of the scan and the identification of additional portions of the physical surface connected to the initial touch location on the top surface of the physical object 7014. In Figure 7M, the scanning and identification of the top surface of the physical object 7014 is complete, and the visual effect expands from the position corresponding to the initial touch location on the top surface of the physical object 7014 to cover the entire top surface of representation 7014'. In some embodiments, the diffusion of the visual effect 7144 stops when the boundaries of the physical surface are identified and the visual effect is applied to the representation of the entire surface. In some embodiments, the visual effect 7144 continues to expand to the representation of additional portions of the physical environment that have been scanned and characterized in the meantime. In some embodiments, the computer system detects the movement of the user's hand 7202, which moves the contact point to another location on the top surface of the physical object 7014 and initiates a new scan from the new touch location of the physical object, or continues the previous scan in parallel with the new scan. In some embodiments, as scanning continues from one or more touch locations, the corresponding visual effect expands from the position corresponding to the touch location based on the scan results. In some embodiments, while the line of sight 7140 is detected on the top surface of the physical object 7014, the computer system detects the user's finger moving across multiple positions along a path on the top surface of the physical object 7014, and optionally performs a scan from a location along the path, expanding the visual effect from the location along the path or area being touched by the user's hand. According to some embodiments, scanning may be performed with higher accuracy and speed than from a single touch point by using depth data at more points on the top surface as a constraint for scanning.
[0180] In Figures 7M to 7N, while displaying a visual effect at a position corresponding to the location of the user's hand touching the top surface of the physical object 7014, according to the physical surface identified by a scan performed by the computer system, the computer system detects the movement of the user's hand that results in the blocking of contact from the top surface of the physical object 7014. Upon detection that the user's hand has left the surface of the physical object 7014, the computer system stops displaying the visual effect at the position of the surface identified based on the scan, as shown in Figure 7N. The representation 7014' is restored to its original appearance before the application of the visual effect 7144 in Figure 7N.
[0181] In some embodiments, after the scan is complete and physical objects and surfaces within a portion of the physical environment have been identified, if the computer system detects user contact with a physical surface (e.g., by the user's hand 7202, another hand, etc.), the computer system optionally redisplays the visual effect 7144 to show the spatial characteristics of the physical surface starting from the position corresponding to the location of the user's touch. In some embodiments, the visual effect is applied to the representation of the entire physical surface as soon as a touch is detected on the physical surface. In some embodiments, the visual effect gradually expands and enlarges across the representation of the physical surface from the position corresponding to the location of the touch.
[0182] In some embodiments, the representation 7014' of the physical object 7014 is provided by a camera view of the physical environment, and the visual effect 7144 replaces the representation of at least a portion of the representation of the physical object 7014' in the view of the three-dimensional environment displayed through the display generating component. In some embodiments, the representation 7014' of the physical object 7014 is provided by a camera view of the physical environment, and the visual effect is projected onto the surface of the physical object, covering a portion of the surface of the physical object in the physical environment, and is seen as part of the camera view of the physical environment. In some embodiments, the representation 7014' of the physical object is part of the view of the physical environment seen through a transparent or translucent portion of the display generating component, and the visual effect is displayed by the display generating component in a position that obstructs the view of at least a portion of the surface of the physical object 7014. In some embodiments, the representation of the physical object 7014' is part of a view of the physical environment seen through the transparent or translucent portion of the display-generating component, and the visual effect is projected onto the surface of the physical object 7014, covering a portion of the surface of the physical object 7014 in the physical environment and being seen as part of the physical environment through the transparent or translucent portion of the display-generating component. In some embodiments, the visual effect is projected directly onto the user's retina, overlaying an image of a portion of the surface of the physical object 7014 on the retina.
[0183] In some embodiments, if the user's hand 7202 touches a different part of the physical environment, such as a wall 7006 or a floor 7008, the computer system applies a visual effect to the position corresponding to the location of the user's touch on the different part of the physical environment or to a surface identified nearby (for example, the visual effect is applied to the vertical plane of the wall representation 7006' or the horizontal plane of the floor representation 7008', etc.).
[0184] In some embodiments, simultaneous detection of gaze and touch input in individual parts of the physical environment is required for the computer system to initiate a scan of a part of the physical environment and / or display visual effects according to the results of the scan of the part of the physical environment. In some embodiments, when the user's gaze is removed from the individual part of the physical environment, the computer system stops displaying the visual effects and optionally stops continuing the scan of the part of the physical environment, even if the user's hand touch is still detected in the individual part of the physical environment.
[0185] In some embodiments, the visual effect 7144 is an animated visual effect that causes an animated visual change in the area to which it is applied. In some embodiments, the animated visual change includes flashing light and / or color changes that change over time in an area within a view of the physical environment to which the visual effect is applied. In some embodiments, when the animated visual change is occurring (e.g., the visual effect affects the appearance of the area using one or more filters or modification functions applied to the original content of the area, but the visual features of the content (e.g., shape, size, type of object, etc.) remain recognizable to the viewer), the area to which the visual effect is applied does not change (e.g., with respect to the size, shape, and / or content displayed in the area, etc.). In some embodiments, the area in a three-dimensional environment to which the visual change is applied expands when the animated visual change is occurring.
[0186] In some embodiments, the computer system applies different visual effects to different parts of a surface that the user's hand touches. In some embodiments, the surface touched by the user's hand extends to an extended region, and the surface properties may differ for different parts of the extended region. In some embodiments, when the user touches the peripheral portion of the extended surface, the visual effect shows an animated movement toward the central portion of the surface representation, and when the user touches the central portion of the extended surface, the visual effect shows a different animated movement toward the peripheral region of the surface representation. In some embodiments, when different visual effects are applied to the same extended region on the surface, the visual effects appear different due to the different starting locations and propagation directions of the animated movement. In some embodiments, different visual effects are generated according to the same baseline visual effect (e.g., gray overlay, flashing visual effect, ripple, expanding mesh wire, etc.), and the difference between different visual effects includes different animations generated according to the same baseline visual effect (e.g., baseline expanding gray overlay with different shaped boundaries, flashing visual effect of baseline modified using different spatial relationships between a virtual light source and a surface below, baseline ripple modified with different wavelengths and / or sources, baseline mesh wire pattern modified with different starting locations, etc.).
[0187] In some embodiments, after the scan is complete and surfaces in the physical environment are identified, the surfaces can be highlighted or visually represented in a view of the physical environment. When the computer system detects contact between the user's hand and a surface that has already been scanned and characterized based on the scan, the computer system displays an animated visual effect that begins at the position on the representation of the surface corresponding to the location of the touch and propagates across the representation of the surface according to the spatial properties of the surface determined based on the scan. In some embodiments, the animated visual effect persists as long as the contact is maintained on the surface. In some embodiments, the computer system requires that the location of the contact remain substantially stationary (e.g., having less than a threshold amount of movement within a threshold time, or not moving at all) in order to continue displaying the animated visual effect. In some embodiments, the computer system requires that the location of the contact remain on the same extended surface (e.g., stationary, or moving within an extended surface) in order to continue displaying the animated visual effect. In some embodiments, the computer system stops displaying the animated visual effect in response to the detection of contact movement across the surface or hand movement away from the surface. In some embodiments, the computer system stops displaying the animated visual effect in response to detecting the movement of the user's hand away from the surface and no longer in contact with the surface. In some embodiments, the computer system stops the animated visual effect in response to detecting contact movement across the surface and / or hand movement away from the surface, and maintains the display of the static state of the visual effect. In some embodiments, the computer system stops the animated visual effect in response to detecting the movement of the user's hand away from the surface and no longer in contact with the surface, and maintains the display of the static state of the visual effect.
[0188] In some embodiments, the visual effects described herein are displayed during the process of generating a spatial representation of at least a portion of the physical environment, and optionally, after the spatial representation of a portion of the physical environment has been generated, in response to the detection of a user's hand touching a portion of the physical environment.
[0189] In some embodiments, the display of the visual effects described herein is triggered when the computer system switches from displaying a virtual reality environment to displaying a representation of the physical environment and / or an augmented reality environment. In some embodiments, the display of the visual effects described herein is triggered when the computer system detects that a display generating component has been positioned in a spatial relationship to the user that allows the user to view the physical environment through the display generating component (e.g., when the HMD is placed on the user's head, in front of the user's eyes, held in front of the user's face, when the user walks or sits in front of the head-up display, when the user turns on the display generating component to view a pass-through view of the physical environment, etc.). In some embodiments, the display of the visual effects described herein is optionally triggered when the computer system switches from displaying a virtual reality environment to displaying a representation of the physical environment and / or an augmented reality environment, without requiring the user to touch any part of the physical environment (e.g., the visual effect is displayed in response to the detection of a line of sight over a part of the physical environment, or optionally initiated at a default location without the user's line of sight, etc.). In some embodiments, the display of the visual effects described herein is triggered when the computer system detects that the display generating component is positioned in a spatial relationship to the user that allows the user to view the physical environment through the display generating component without requiring the user to touch any part of the physical environment (for example, the visual effect is displayed in response to the detection of a line of sight over a part of the physical environment, or is optionally initiated at a default location without the user's line of sight).
[0190] Figures 7O to 7Q are block diagrams illustrating how, in several embodiments, an interactive user interface object is displayed at a position in a three-dimensional environment corresponding to a first part of the physical environment (e.g., the location of a physical surface, a location in free space within the physical environment, etc.), and how the display of individual sub-parts of the user interface object can be selectively deselected according to the location of a part of the user (e.g., the user's fingers, hands, etc.) moving in space between the first part of the physical environment and the location corresponding to the viewpoint of the currently displayed view of the three-dimensional environment.
[0191] In some embodiments, the computer system displays interactive user interface objects (e.g., user interface object 7152, another user interface object such as a control panel, a user interface object containing selectable options, an integrated control object, etc.) in a three-dimensional environment (e.g., environment 7151, or another environment, etc.). The computer system also displays a representation of a physical environment within the three-dimensional environment (e.g., environment 105 in Figure 7I, another physical environment, etc.), and the interactive user interface objects have individual spatial relationships to various positions in the three-dimensional environment corresponding to different locations in the physical environment. When a user interacts with a three-dimensional environment with one or more fingers or a part of the user's hand (e.g., hand 7202, fingers of hand 7202, etc.) through touch input and / or gesture input, the user's part (e.g., a part of the user's hand, the whole hand, and possibly the wrist and arm connected to the hand, etc.) can enter a spatial region between locations corresponding to the positions of user interface objects (e.g., locations of physical objects or physical surfaces, locations in free space within the physical environment, etc.) and locations corresponding to the viewpoint of the currently displayed view of the three-dimensional environment (e.g., the location of the user's eyes, locations of display-generating components, locations of cameras capturing the view of the physical environment shown in the three-dimensional environment, etc.). Based on the spatial relationships between the user's hand location, the locations corresponding to the positions of user interface objects, and the locations corresponding to the viewpoint, the computer system determines which parts of the user interface objects are visually blocked by the user's part and which parts of the user interface objects are not visually blocked by the user's part when viewed by the user from the viewpoint location.Next, the computer system stops displaying individual parts of the user interface object that would be visually blocked by the user's part (as determined by the computer system, for example), and instead, as shown in Figure 7P, maintains the display of other parts of the user interface object that would not be visually blocked by the user's part (as determined by the computer system, for example), while allowing the representation of the user's part to be visible at the position of the individual parts of the user interface object. In some embodiments, in response to the detection of movement of the user's part or viewpoint (for example, due to movement of display generation components, movement of a camera capturing the physical environment, movement of the user's head or torso, etc.), the computer system re-evaluates which parts of the user interface object are visually blocked by the user's part and which parts are not visually blocked by the user's part when viewed by the user from the viewpoint location, based on the new spatial relationships between the user's part, the location corresponding to the viewpoint, and the location corresponding to the position of the user interface object. The computer system then stops displaying other parts of the user interface object that would be visually blocked by the user's part (as determined by the computer system, for example), and allows the previously stopped-from-display parts of the user interface object to be restored to the view of the three-dimensional environment, as shown in Figure 7Q.
[0192] In some embodiments, when a user interacts with user interface objects (e.g., user interface object 7152, another user interface object such as a control panel, a user interface object with selectable options, an integrated control object, etc.) in an augmented reality or virtual reality environment, the tactile sensations provided by the physical surfaces of the physical environment help to better orient the user's spatial sense in the augmented reality or virtual reality environment, and as a result, the user can provide more precise input when interacting with user interface objects. In some embodiments, the physical surface may include touch sensors that provide more precise information about the user's touch on the physical surface (e.g., touch location, touch duration, touch intensity, etc.), which enables more diverse and / or sophisticated input for interacting with user interface objects or parts thereof. In some embodiments, the physical surface may include surface features (e.g., bumps, buttons, textures, etc.) that help the user precisely position their gestures or touch inputs relative to the surface features, and may also provide a more realistic experience of interacting with user interface objects that have visual features (e.g., virtual markers, buttons, textures, etc.) corresponding to the surface features on the physical surface.
[0193] As described herein, when a user interface object is displayed in a position corresponding to the location of a physical surface having spatial properties corresponding to the spatial properties of the physical surface, the user interface object appears to overlay or extend the representation of the physical surface or a virtual surface having the spatial properties of the physical surface. To provide the user with a more realistic and intuitive experience when the user interacts with the user interface object via touch input on a physical surface, the user interface object is visually segmented into multiple parts, and when the user's hand is in a portion of the physical space between the individual part of the physical surface and the user's eyes, at least one of the multiple parts is visually obscured by the representation of the user's hand. In other words, at least a portion of the user's hand (optionally, other parts of the user connected to the hand) may intersect with the user's line of sight directed towards the individual part of the user interface object, obstructing the user's view of the individual part of the user interface object. In some embodiments, as the user's hand moves within the space between the physical surface and the user's eyes, at least a portion of the user's hand (optionally, other parts of the user connected to the hand) may intersect the user's line of sight directed to different parts of the user interface object, blocking the user's view of those parts and allowing previously blocked portions of the user interface object to be revealed again.
[0194] In some embodiments, a physical surface includes one or more parts having spatial contours and surface textures corresponding to different types of user interface elements such as buttons, sliders, bumps, circles, checkmarks, and switches. In some embodiments, individual parts of a user interface object corresponding to individual user interface elements are optionally segmented into multiple sub-parts, where only a portion of the sub-parts is visually obscured by the representation of the user's hand in the view of the three-dimensional environment, while a portion of the sub-parts of the user interface element is not visually obscured by the representation of the user's hand in the view of the three-dimensional environment.
[0195] In Figures 7O to 7Q, the display generation component 7100 displays a view of the three-dimensional environment 7151. In some embodiments, the three-dimensional environment 7151 is a virtual three-dimensional environment that includes virtual objects and virtual surfaces at various spatial positions within the three-dimensional environment. In some embodiments, the three-dimensional environment 7151 is an augmented reality environment that includes a representation of a physical environment having representations of physical objects and surfaces placed at various positions corresponding to each location in the physical environment, and virtual content having positions relative to the positions of the representations of physical objects and surfaces in the three-dimensional environment. In some embodiments, the view of the three-dimensional environment includes at least a first surface (e.g., a virtual surface, or a representation of a physical surface) at a position corresponding to the location of the first physical surface, and has spatial properties (e.g., orientation, size, shape, surface profile, surface texture, spatial extent, etc.) corresponding to the spatial properties (e.g., orientation, size, shape, surface profile, surface texture, spatial extent, etc.) of the first physical surface in the physical environment. In some embodiments, physical surfaces include tabletop surfaces, wall surfaces, display device surfaces, touchpad surfaces, user's lap surfaces, palm surfaces, buttons, and the surfaces of prototype objects with hardware affordances. In this example, the top surface of physical object 7014 is used as a non-limiting example of a physical surface that the user's hand touches.
[0196] In this example, a first user interface object (e.g., a virtual keyboard 7152, a control panel with one or more control affordances, a menu with selectable options, a single integrated control object, etc.) containing one or more interaction parts corresponding to each action is displayed in a location within the three-dimensional environment 7151 that corresponds to the position of a first physical surface (e.g., the top surface of a physical object 7014 represented by representation 7014', the surface of a physical object at a location corresponding to the position of the virtual object 7014', etc.). The spatial properties of the first user interface object (e.g., a virtual keyboard 7152, a control panel with one or more control affordances, a menu with selectable options, a single integrated control object, etc.) correspond to the spatial properties of the first physical surface. For example, if the first user interface object is planar and the first physical surface is planar, it will be displayed parallel to the representation of the first physical surface. In another example, in some embodiments, the first user interface object has a surface profile corresponding to the surface profile of the first physical surface, and the positions of topological features (e.g., bumps, buttons, textures, etc.) on the first user interface object are aligned with positions corresponding to the locations of the corresponding topological features on the first physical surface. In some embodiments, the first user interface object has topological features that do not exist at locations on the first physical surface corresponding to the positions of the topological features on the first user interface object.
[0197] As shown in Figure 7O, the computer system displays a view of the three-dimensional environment 7151 via the display generation component 7100. According to some embodiments, the view of the three-dimensional environment 7151 includes physical surfaces (e.g., representations of vertical walls 7004 and 7006 7004' and 7006', representation of horizontal floor 7008 7008, representation of the surface of a physical object, etc.) and object representations (e.g., representation of physical object 7014 7014', representation of other physical objects, etc.) in a portion of the physical environment as seen from a viewpoint corresponding to the location of the display generation component 7100 in the physical environment (e.g., the location corresponding to the user's eyes or head if the display generation component 7100 is an HMD). According to some embodiments, the spatial relationship between physical surfaces and physical objects in the physical environment 105 is represented in the three-dimensional environment by the spatial relationship between the representations of physical surfaces and physical objects in the three-dimensional environment 7151. In some embodiments, when a user moves a display generation component relative to the physical environment, the viewpoint of the currently displayed view moves within the three-dimensional environment, resulting in different views of the three-dimensional environment 7151 from different viewpoints. In some embodiments, a computer system performs a scan of the environment to identify surfaces and planes and constructs a three-dimensional model of the physical environment. According to some embodiments, the computer system defines the position of a virtual object relative to the three-dimensional model, and as a result, the virtual object can be positioned within the three-dimensional environment in various spatial relationships with respect to physical surfaces and object representations within the three-dimensional environment. For example, a virtual object may optionally be given an upright orientation relative to the three-dimensional environment 7151 and displayed in a position and / or orientation that simulates a specific spatial relationship (e.g., overlapping, standing, parallel, perpendicular, etc.) with a physical surface or object representation (e.g., representation 7014' of physical object 7014, representation 7008' of floor 7008, etc.).
[0198] In Figures 7O to 7Q, the first user interface object (e.g., the virtual keyboard 7152 in this example) is displayed at a location corresponding to the location of the first physical surface (e.g., the top surface of the physical object 7014 represented by representation 7014', or the top surface of a physical object located in a position corresponding to the top surface of the virtual object 7014'), and the spatial properties of the first user interface object correspond to the spatial properties of the first physical surface (e.g., being parallel to the first physical surface, conforming to the surface profile of the first physical surface, etc.). In some embodiments, the computer system moves the first user interface object in response to the movement of the first physical surface in the physical environment. For example, in some embodiments, the first user interface object remains displayed in the same spatial relationship as the representation of the first physical surface in the three-dimensional environment while the first physical surface is moving in the physical environment.
[0199] In some embodiments, the representation 7014' of the physical object 7014 is provided by a camera view of the physical environment, and the first user interface object replaces the display of at least a portion of the representation 7104' of the physical object in a view of a three-dimensional environment (e.g., environment 7151, or another augmented reality environment) displayed through the display generation component. In some embodiments, the representation 7014' of the physical object is provided by a camera view of the physical environment, and the first user interface object is projected onto the surface of the physical object 7014, covering a portion of the surface of the physical object 7014 in the physical environment, and is seen as part of the camera view of the physical environment. In some embodiments, the representation 7014' of the physical object 7014 is part of the view of the physical environment seen through a transparent or translucent portion of the display generation component, and the first user interface object is displayed by the display generation component in a position that obstructs the view of at least a portion of the representation 7014' of the physical object 7014'. In some embodiments, the representation 7014' of the physical object 7014 is part of a view of the physical environment seen through the transparent or translucent portion of the display-generating component, and the first user interface object is projected onto the surface of the physical object 7014, covering part of the surface of the physical object 7014 in the physical environment, and is seen as part of the physical environment through the transparent or translucent portion of the display-generating component. In some embodiments, the first user interface object is an image projected onto the user's retina, overlaying part of the image of the surface of the physical object 7014 onto the user's retina (for example, the image is an image of a camera view of the physical environment provided by the display-generating component, or an image of a view of the physical environment through the transparent portion of the display-generating component).
[0200] In the examples of Figures 7O to 7Q, the first user interface object (e.g., the virtual keyboard 7152 in this example) is not visually obscured by the representation of the user's hand 7202' in the view of the three-dimensional environment 7151 before the user interacts with the first user interface object. In some embodiments, the representation of the user's hand 7202' may be visually obscured by the presence of other user interface objects (e.g., a text input window 7150, or other user interface objects) depending on the spatial relationship between the positions of the other user interface objects and the positions corresponding to the location of the user's hand 7202 (e.g., the position of the representation of the user's hand 7202 7202' in the three-dimensional environment). For example, because the virtual position of the representation of the user's hand 7202 7202' is further from the viewpoint of the currently displayed view of the environment 7151 than the text input window 7150, which is aligned with the user's line of sight in the environment 7151, a portion of the representation of the hand 7202 7202' is blocked by the text input box 7150 in Figure 7O. In some embodiments, the representation 7202' of the user's hand 7202 is part of a camera view of the physical environment. In some embodiments, the representation 7202' of the user's hand 7202 is a view of the hand through a transparent portion of a display generation component. In some embodiments, the representation of the user's hand is a stylized representation created based on real-time data regarding the shape and location of the hand in the physical environment.
[0201] In Figure 7P, the user's hand 7202 moves toward a first physical surface in the physical environment (for example, the top surface of the physical object 7014 represented by representation 7014' in this example). In some embodiments, a part of the hand 7202, such as one or more fingers (e.g., the index finger, thumb, and middle finger), makes contact with the first physical surface at a first location on the first physical surface. In some embodiments, the first location on the first physical surface corresponds to a first position on a first user interface object, and the first position on the first user interface object corresponds to a first action associated with the first user interface object. In this particular example, the first location on the first physical surface corresponds to the position of the character key "I" (e.g., key 7154 in this example) on the virtual keyboard 7152, and the first action associated with the first user interface object is to enter the text character "I" (e.g., character 7156 in this example) into the text input window 7150. In some embodiments, the first user interface object is a control panel, the first location on the first physical surface corresponds to the position of a first control object (e.g., a button, slider, switch, checkbox, etc.) within the first user interface object, and the first operation associated with the first user interface object is an operation associated with the first control object, such as turning a device or function on or off, adjusting the value of a control function, or selecting a parameter of a function or setting. When contact of the user's hand 7202 is detected at the first location on the first physical surface, the computer system identifies the corresponding control object on the first user interface object, performs the first operation, and optionally updates the appearance of the first control object and / or the environment 7151 to indicate that the first operation has been performed.
[0202] In some embodiments, the computer system considers the relationship between the first physical surface and the user's hand (e.g., one or more fingers of the user's hand), the shape (e.g., circular, elongated, etc.), size (e.g., small, large, etc.), duration (e.g., less than the threshold duration for a tap input, longer than the threshold duration for a long tap input, longer than the threshold duration for a touch-hold input without lift-off, etc.), direction of movement (e.g., up, down, left, right, clockwise, counterclockwise, etc.), and distance of movement (e.g., For example, the computer system determines the characteristics of the contact between the user interface object and the first position in the first user interface object, such as: less than a threshold amount of movement within a threshold time, greater than a threshold amount of movement within a threshold time, greater than a threshold amount of translation, greater than a threshold amount of rotation, etc.; the movement path (e.g., straight path, curved path, zigzag path, intersecting a threshold position / angle, not intersecting a threshold position / angle, etc.); contact strength (e.g., greater than threshold strength, less than threshold strength, etc.); number of contacts (e.g., single contact, two contacts, etc.); and repetition of repeated contacts (e.g., double tap, triple tap, etc.), and two or more combinations of the above. Based on the characteristics of the contact, the computer system determines which of a plurality of actions associated with the first user interface object and / or the first position in the first user interface object should be performed. In some embodiments, the computer system evaluates the contact against various predefined criteria and, according to the determination that the contact satisfies the predefined criteria corresponding to the individual action (e.g., regardless of the characteristics of the contact (e.g., starting an experience, turning a function on / off, etc.), according to the characteristics of the contact (e.g., adjusting a value, performing a series of actions with an adjustable parameter, etc.), etc.).
[0203] In some embodiments, as shown in Figure 7P, while the user's hand is in a spatial region within the physical environment between a location corresponding to the viewpoint position (e.g., the location of a display generating component, the location of the user's eyes, the current view of the user's hand, and the location of a camera capturing the physical environment shown in the view of the environment 7151) and the first physical surface, the computer system maintains the display of the second part of the first user interface object while deactivating or stopping the display of the first part of the first user interface object, so that the user's hand appears in the view of the three-dimensional environment 7151 at the position of the first part of the first user interface object. For example, as shown in Figure 7P, a first portion of the virtual keyboard 7152 that is located at a location corresponding to the location behind the user's hand 7020 relative to a position corresponding to the viewpoint (e.g., the location of the display generation component, the user's eyes, the computer system's camera, etc.) (e.g., a portion of key 7154, portions of the two keys directly above key 7154, and portions of the top two keys above key 7154, etc.) is not displayed in the view of the three-dimensional environment 7151, while the rest of the virtual keyboard 7152 that is not behind the location of the user's hand 7020 continues to be displayed in the view of the three-dimensional environment 7151. In some embodiments, a portion of the user's hand that is not in contact with the first physical surface may enter a spatial region within the physical environment between the first physical surface and a location corresponding to the viewpoint position (e.g., the location of a display generating component, the location of the user's eyes, the current view of the user's hand, and the location of a camera capturing the physical environment shown in the view of the environment 7151), and the computer system may remove or stop displaying any portion of the first user interface object that is at a position corresponding to a location that is visually blocked by the portion of the user's hand when viewed from the location corresponding to the current viewpoint of the three-dimensional environment. For example, any portion of the keys at a position in the virtual keyboard 7152 behind the location of the user's thumb relative to the viewpoint location may also not be displayed.In Figure 7P, the position of the text input window 7150 is in front of the position corresponding to the location of the user's hand 7202, so the text input window 7150 is displayed in the view of the three-dimensional environment 7151, blocking (or replacing) part of the view of the representation 7202' of the user's hand 7202.
[0204] Figure 7Q shows that while displaying a view of the three-dimensional environment 7151, the computer system detects the user's hand movements within the physical environment. For example, the movements include being lifted from a first location on a first physical surface of a physical object and moving to another location on the first physical surface of the physical object. In some embodiments, a part of the hand 7202, such as one or more fingers (e.g., index finger, thumb, and middle finger), makes contact with the first physical surface at a second location on the first physical surface. In some embodiments, the second location on the first physical surface corresponds to a second position on a first user interface object, and the second position on the first user interface object corresponds to a second action associated with the first user interface object. In this particular example, the second location on the first physical surface corresponds to the position of the character key "p" (e.g., key 7160 in this example) on the virtual keyboard 7152, and the second action associated with the first user interface object is to enter the text character "p" (e.g., character 7158 in this example) into the text input window 7150. In some embodiments, the first user interface object is a control panel, the second location on the first physical surface corresponds to the position of a second control object (e.g., a button, slider, switch, checkbox, etc.) within the first user interface object, and the second action associated with the first user interface object is an action associated with the second control object, such as turning a device or function on or off, adjusting the value of a control function, or selecting a parameter of a function or setting. When user hand contact is detected at the second location, the computer system identifies the corresponding control object on the first user interface object, performs the second action, and optionally updates the appearance of the second control object and / or environment 7151 to indicate that the second action has been performed.In some embodiments, the computer system considers the relationship between the first physical surface and the user's hand (e.g., one or more fingers of the user's hand), the shape (e.g., circular, elongated, etc.), size (e.g., small, large, etc.), duration (e.g., less than the threshold duration for a tap input, longer than the threshold duration for a long tap input, longer than the threshold duration for a touch-hold input without lift-off, etc.), direction of movement (e.g., up, down, left, right, clockwise, counterclockwise, etc.), and distance of movement (e.g., For example, the computer system determines the characteristics of the contact between the object and the contact (e.g., less than the threshold amount of movement within a threshold time, greater than the threshold amount of movement within a threshold time, greater than the threshold amount of translation, greater than the threshold amount of rotation, etc.), the movement path (e.g., straight path, curved path, zigzag path, intersecting the threshold position / angle, not intersecting the threshold position / angle, etc.), the contact strength (e.g., greater than the threshold strength, less than the threshold strength, etc.), the number of contacts (e.g., single contact, two contacts, etc.), the repetition of repeated contacts (e.g., double tap, triple tap, etc.), and two or more combinations of the above. Based on the characteristics of the contact, the computer system determines which of a plurality of actions associated with the first user interface object and / or the second position of the first user interface object should be performed. In some embodiments, the computer system evaluates the contact against various predefined criteria and, according to the determination that the predefined criteria corresponding to the individual action are met by the contact, the computer system performs the individual action (e.g., regardless of the characteristics of the contact, according to the characteristics of the contact, etc.).
[0205] In some embodiments, as shown in Figure 7Q, while the user's hand 7202 is in a spatial region within the physical environment between the location corresponding to the viewpoint position (e.g., the location of the display generation component, the location of the user's eyes, the current view of the user's hand, and the location of the camera that captures the physical environment shown in the view of the environment 7151) and the first physical surface, the computer system maintains the display of the fourth part of the first user interface object while deactivating or stopping the display of the third part of the first user interface object, so that the user's hand appears in the view of the three-dimensional environment 7151 at the position of the third part of the first user interface object. For example, as shown in Figure 7Q, a third portion of the virtual keyboard 7152 that is located at a location corresponding to the location behind the user's hand 7202 relative to a viewpoint (e.g., a portion of key 7160, portions of the two keys directly above key 7160, and portions of the top two keys above key 7160) is not displayed in the view of the three-dimensional environment 7151, while the other portion of the virtual keyboard 7152 that is not behind the location of the user's hand 7020 continues to be displayed in the view of the three-dimensional environment 7151. In Figure 7Q, since the location corresponding to the location of the user's hand 7202 is no longer behind the position of the text input window 7150, the text input window 7150 displayed in the view of the three-dimensional environment 7151 no longer blocks the view of the representation of the user's hand 7202' (or replaces the display of a portion of the hand representation 7202'). As shown in Figure 7Q, the first part of the virtual keyboard 7152 (e.g., key 7154, the two keys above key 7154, etc.), which was previously hidden by the presence of the hand 7202' representation 7202', is no longer hidden and is again visible in the view of the three-dimensional environment 7151.
[0206] In some embodiments, the first user interface object is a single user interface object, such as a single button, a single checkbox, or a single selectable option, and a pre-configured user input detected at a first, second, or third location on a first physical surface causes the computer system to perform the same action associated with the first user interface object, where the first, second, and third locations correspond to the first, second, and third parts of the single user interface object, respectively. In some embodiments, depending on the location of the user's hand in the physical environment, the computer system selectively stops displaying one individual part of the first, second, or third part of the integrated user interface object based on the determination that the user's hand is between the viewpoint location and the location of the user's hand in the physical environment.
[0207] In some embodiments, there are multiple user interface objects displayed at positions in a three-dimensional environment 7151 corresponding to different positions in the physical environment, and when the user's hand is present in the spatial portion of the physical environment between the viewpoint location and the locations corresponding to different user interface object locations, the computer system segments the multiple user interface objects and selectively stops displaying each portion of the multiple user interface objects that has a position corresponding to a location blocked by the presence of the user's hand when viewed from the location corresponding to the current viewpoint in the three-dimensional environment 7151. In some embodiments, even if the representation of the user's hand simultaneously removes portions of both the first and second user interface objects from the view of the three-dimensional environment, the user's hand interacts with the first user interface object and does not activate the second user interface object in the same view of the three-dimensional environment. For example, in Figures 7P and 7Q, even if the computer system stops displaying portions of multiple keys on a virtual keyboard, only the keys at positions corresponding to the user's touch location or the location of a specific part of the user's hand (e.g., the tip of the index finger, the tip of the thumb, etc.) are activated.
[0208] In some embodiments, the computer system determines, for example, the shape and position of the user's hand and the location of a virtual light source in the three-dimensional environment, the shape and position of the simulated shadow of the representation of the user's hand 7202 in the view of the three-dimensional environment 7151. The computer system displays the simulated shadow on the position on the surface of the first user interface object by optionally changing the appearance of the portion of the first user interface object at that position, or by replacing the display of the portion of the first user interface object at that position.
[0209] In some embodiments, the input gestures used in the various examples and embodiments described herein (for example, relating to Figures 7A–7Q and Figures 8–11) optionally include discrete small motion gestures performed by moving one or more of the user's fingers relative to other fingers or parts of the user's hand, without optionally requiring the user to move their entire hand or arm significantly away from their natural location(s) and orientation(s) to perform an action immediately before or during the gesture in order to interact with a virtual or mixed reality environment.
[0210] In some embodiments, input gestures are detected by analyzing data and signals captured by a sensor system (e.g., sensor 190 in Figure 1, image sensor 314 in Figure 3). In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras such as a motion RGB camera, an infrared camera, a depth camera, etc.). For example, one or more imaging sensors are components of a computer system (e.g., computer system 101 in Figure 1 (e.g., portable electronic device 7100 or HMD)) that includes display generation components (e.g., display generation component 120 in Figures 1, 3, and 4 (e.g., a touchscreen display that functions as a display and a touch-sensing surface, a stereoscopic display, a display with a pass-through portion, etc.)) or provide data to the computer system. In some embodiments, one or more imaging sensors include one or more rear cameras on the side of the device opposite to the device's display. In some embodiments, input gestures are detected by a sensor system of a head-mounted system (e.g., a VR headset that includes a stereoscopic display that provides a left image for the user's left eye and a right image for the user's right eye). For example, one or more cameras, which are components of a head-mounted system, are mounted on the front and / or bottom of the head-mounted system. In some embodiments, one or more imaging sensors are positioned in the space in which the head-mounted system is used (e.g., arranged around the head-mounted system at various locations in a room) so that the imaging sensors capture images of the head-mounted system and / or the user of the head-mounted system. In some embodiments, input gestures are detected by a sensor system of a head-up device (e.g., a head-up display, a car windshield capable of displaying graphics, a window capable of displaying graphics, or a lens capable of displaying graphics). For example, one or more imaging sensors are mounted on the interior of a car. In some embodiments, the sensor system includes one or more depth sensors (e.g., a sensor array).For example, one or more depth sensors include one or more light-based (e.g., infrared) sensors and / or one or more acoustic-based (e.g., ultrasonic) sensors. In some embodiments, the sensor system includes one or more signal emitters, such as light emitters (e.g., infrared emitters) and / or sound emitters (e.g., ultrasonic emitters). For example, while light (e.g., light from an infrared light emitter array having a predetermined pattern) is projected onto a hand (e.g., hand 7200), an image of the hand under illumination is captured by one or more cameras, and the captured image is analyzed to determine the position and / or configuration of the hand. In contrast to using signals from a touch-sensing surface or other direct contact or proximity-based mechanisms, determining input gestures using signals from image sensors directed at the hand allows the user to freely choose whether to perform large movements or remain relatively stationary when providing input gestures with their hand, without experiencing constraints imposed by a particular input device or input area.
[0211] In some embodiments, a tap input optionally indicates a tap of the thumb on the index finger of the user's hand (e.g., on the side of the index finger adjacent to the thumb). In some embodiments, a tap input is detected without requiring the thumb to be lifted from the side of the index finger. In some embodiments, a tap input is detected according to a determination that a downward movement of the thumb is followed by an upward movement of the thumb, and the thumb is in contact with the side of the index finger for less than a threshold time. In some embodiments, a tap-hold input is detected according to a determination that the thumb moves from an elevated position to a touch-down position and remains in the touch-down position for at least a first threshold time (e.g., a tap-time threshold or another time threshold longer than the tap-time threshold). In some embodiments, the computer system requires the entire hand to remain substantially stationary at a location for at least a first threshold time in order to detect a tap-hold input by the thumb on the index finger. In some embodiments, a touch-hold input is detected without requiring the hand to remain substantially stationary (e.g., the entire hand can move while the thumb is placed on the side of the index finger). In some embodiments, tap-hole drag input is detected when the thumb touches the side of the index finger and the entire hand moves while the thumb remains stationary on the side of the index finger.
[0212] In some embodiments, a flick gesture optionally represents a push or flick input of the thumb moving across the index finger (e.g., from the palmar side to the rear side of the index finger). In some embodiments, an extension of the thumb involves an upward movement away from the side of the index finger, such as an upward flick input by the thumb. In some embodiments, the index finger moves in the opposite direction to the thumb while the thumb moves forward and upward. In some embodiments, a reverse flick input is performed by the thumb moving from an extended position to a retracted position. In some embodiments, the index finger moves in the opposite direction to the thumb while the thumb moves backward and downward.
[0213] In some embodiments, the swipe gesture is optionally a swipe input by moving the thumb along the index finger (e.g., along the side of the index finger adjacent to the thumb or along the side of the palm). In some embodiments, the index finger is optionally extended (e.g., substantially straight) or flexed. In some embodiments, the index finger moves between the extended and flexed states while the thumb moves in the swipe input gesture.
[0214] In some embodiments, different phalanges of different fingers correspond to different inputs. Thumb tap inputs across different phalanges of different fingers (e.g., index finger, middle finger, ring finger, and optionally little finger) are optionally mapped to different actions. Similarly, in some embodiments, different push or click inputs performed by the thumb across different fingers and / or different parts of the fingers can trigger different actions in individual user interface contacts. Likewise, in some embodiments, different swipe inputs performed by the thumb along different fingers and / or in different directions (e.g., towards the distal or proximal end of the finger) trigger different actions in their respective user interface contexts.
[0215] In some embodiments, the computer system processes tap input, flick input, and swipe input as different types of input based on the type of thumb movement. In some embodiments, the computer system processes input having different finger locations tapped, touched, or swiped by the thumb as different sub-input types (e.g., proximal, intermediate, distal subtypes, or index finger, middle finger, ring finger, or little finger subtypes) of a given input type (e.g., tap input type, flick input type, swipe input type, etc.). In some embodiments, the amount of movement performed by the moving finger (e.g., thumb), and / or other measures of movement associated with the finger movement (e.g., velocity, initial velocity, ending velocity, duration, direction, movement pattern, etc.) are used to quantitatively influence the action triggered by the finger input.
[0216] In some embodiments, the computer system recognizes combination input types that combine a series of thumb movements, such as tap-swipe input (e.g., the thumb swiping along the side of a finger after touching down another finger), tap-flick input (e.g., the thumb flicking across a finger from the side of the palm to the back of the finger after touching down another finger), and double-tap input (e.g., two consecutive taps on the side of a finger at approximately the same location).
[0217] In some embodiments, gesture input is performed by the index finger instead of the thumb (e.g., the index finger performs a tap or swipe on the thumb, or the thumb and index finger move toward each other to perform a pinch gesture). In some embodiments, wrist movement (e.g., a wrist flick in a horizontal or vertical direction) is performed immediately before, immediately after (e.g., within a threshold time), or concurrently with finger movement input to trigger an additional, different, or modified action in the current user interface context compared to finger movement input without modification by wrist movement. In some embodiments, finger input gestures performed with the user's palm facing the user's face are treated as a different type of gesture than finger input gestures performed with the user's palm facing away from the user's face. For example, a tap gesture performed with the user's palm facing the user performs an action with added (or reduced) privacy protection compared to an action performed in response to a tap gesture performed with the user's palm facing away from the user's face (e.g., the same action).
[0218] In the embodiments provided herein, one type of finger input can be used to trigger an action type, but in other embodiments, other types of finger input may be optionally used to trigger the same type of action.
[0219] Further explanations regarding Figures 7A to 7Q are provided below with reference to methods 8000, 9000, 10000, and 11000 described with respect to Figures 8 to 11.
[0220] Figure 8 is a flowchart of a method 8000 for selecting different audio output modes depending on the level of immersion to which computer-generated content is presented, according to several embodiments.
[0221] In some embodiments, Method 8000 is performed on a computer system (e.g., computer system 101 in Figure 1). It includes display generating components (e.g., display generating components 120 in Figures 1, 3, and 4) (e.g., head-up display, display, touchscreen, projector, etc.) and one or more cameras (e.g., cameras (e.g., color sensor, infrared sensor, and other depth-sensing cameras)) which are cameras facing forward from the user's hand or the user's head. In some embodiments, Method 8000 is stored on a non-temporary computer-readable storage medium and is performed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 in Figure 1A). Some operations of Method 8000 are optionally combined, and / or the order of some operations is optionally changed.
[0222] In some embodiments, Method 8000 is performed by a computer system (e.g., computer system 101 in Figure 1) that communicates with a first display generation component (e.g., display generation component 120, display generation component 7100, etc., in Figures 1, 3, and 4) (e.g., a head-up display, HMD, display, touchscreen, projector, etc.), one or more audio output devices (e.g., earphones, speakers placed in the physical environment, speakers in the same housing, or speakers mounted on the same support structure as the first display generation component (e.g., built-in speakers of the HMD, etc.)), and one or more input devices (e.g., a camera, controller, touch-sensing surface, joystick, button, glove, watch, motion sensor, compass sensor, etc.). In some embodiments, the first display generation component is a user-facing display component that provides the user with a CGR experience. In some embodiments, the computer system is an integrated device having one or more processors and memory enclosed in the same housing as the first display generation component, one or more audio output devices, and at least some of the one or more input devices. In some embodiments, the computer system includes a computing component (e.g., a server, a mobile electronic device such as a smartphone or tablet device, a wearable device such as a watch, a wristband or earphone, a desktop computer, a laptop computer, etc.) which includes one or more processors and a memory separate from one or more of the one or more input devices, which include one or more display generating components (e.g., a head-up display, a touchscreen, a standalone display, etc.), one or more output devices (e.g., earphones, an external speaker, etc.). In some embodiments, the display generating components and one or more audio output devices are integrated and enclosed within the same housing.
[0223] Method 8000 involves a computer system displaying a three-dimensional computer-generated environment (e.g., environment 7102 in Figures 7A-7B, or another three-dimensional environment) via a first display generation component (8002) (for example, displaying a three-dimensional computer-generated environment includes displaying a three-dimensional virtual environment, a three-dimensional augmented reality environment, a pass-through view of a physical environment having a corresponding computer-generated three-dimensional model corresponding to the spatial characteristics of the physical environment, etc.). While displaying a three-dimensional computer-generated environment, the computer system detects a first event corresponding to a request to present a first computer-generated content (8004) (e.g., detecting user input to select and / or activate an icon corresponding to the first computer-generated content, detecting trigger conditions to start the first computer-generated content that are met by user actions or other internal events of the computer system), and the first computer-generated content includes first visual content (e.g., video content, game content, animation, user interface, movie, etc.) and first audio content corresponding to the first visual content (e.g., video content and associated audio data, with timing data relating different parts of the video content to different parts of the audio data (e.g., the video playback timeline and the audio playback timeline are temporally correlated by the timing data)) (e.g., sound effects, soundtrack, audio recording, movie soundtrack, game soundtrack, etc.). For example, the first computer-generated content includes the first visual content 7106 in Figures 7A-7B.In response to the detection of a first event (8006) corresponding to a request to present first computer-generated content, the first event is a first immersion level, and the first computer-generated content presented at the first immersion level occupies a first portion of the three-dimensional computer-generated environment, according to the determination that it corresponds to an individual request to present first computer-generated content having a first immersion level (e.g., an intermediate level of immersion among several available immersion levels, the lowest level of immersion among two or more available immersion levels, or a lower level of immersion among two or more available immersion levels), (e.g., playing video content in a window that occupies a portion of the user's field of view for the three-dimensional computer-generated environment, the current state of the three-dimensional computer-generated environment, where the three-dimensional computer-generated environment extends beyond a preset threshold angle from the viewpoint) (Playing video content having a field of view extending below a preset threshold angle in a three-dimensional computer-generated environment from a viewpoint corresponding to the view), the computer system displays first visual content in a first part of the three-dimensional environment (8008) (e.g., optionally, simultaneously with representations of other virtual content and / or other parts of the physical environment occupying other parts of the three-dimensional computer-generated environment), and the computer system outputs first audio content using a first audio output mode (e.g., stereo audio mode, surround sound mode, etc.) (e.g., minimum immersion audio output mode from several available audio output modes for the first audio content, intermediate level immersion audio mode from several available audio output modes for the first audio content, less immersion audio output mode from several available audio output modes for the first audio content, etc.).In response to the detection of a first event corresponding to a request to present first computer-generated content (8006), and in accordance with the determination that the first event corresponds to a specific request to present first computer-generated content at a second immersion level different from the first immersion level, and that the first computer-generated content presented at the second immersion level occupies a second portion of the three-dimensional computer-generated environment that is larger than the first portion of the three-dimensional environment (for example, instead of occupying a two-dimensional window in the three-dimensional environment, the display of the content occupies a span of three-dimensional space larger than the window; instead of spanning a part of the three-dimensional environment, the visual content extends to the entire three-dimensional environment), the computer system displays the first visual content in the second portion of the three-dimensional environment (8010) (for example, simultaneously with, optionally, other virtual content and / or representations of the physical environment occupying other parts of the three-dimensional environment), and the computer system displays the first visual content in the second portion of the three-dimensional environment (8010), in a manner different from the first audio output mode. Outputting the first audio content using a second audio output mode (e.g., surround sound mode, spatial audio mode with sound localization based on the location of virtual sound sources in the first computer-generated content) (e.g., a more immersive audio output mode among several available audio output modes for the first audio content, an audio mode with the highest level of immersion among several available audio output modes for the first audio content, the most immersive audio output mode among several available audio output modes for the first audio content), and using the second audio output mode instead of the first audio output mode, modifies the level of immersion of the first audio content (e.g., automatically without requiring user input) (e.g., making the first audio content more or less immersive, more or less spatially expanded, having more or less complex spatial changes, and more or less directionally adjustable based on the corresponding visual content).This is shown in Figures 7A and 7B, where Figure 7A shows the display of computer-generated content 7106 using a first immersion level, and Figure 7B shows the display of computer-generated content 7106 using a second immersion level. The computer-generated content displayed at the first immersion level has a smaller spatial range than the computer-generated content displayed at the second immersion level, and the computer system selects different audio output modes for outputting the audio content of the computer-generated content based on the immersion level at which the computer-generated content is displayed by the display generation component.
[0224] In some embodiments, outputting first audio content using a first audio output mode includes outputting first audio content using a first set of sound sources located in a first set of locations in the physical environment (e.g., two sound source outputs located on either side of the HMD, a single sound source located in front of the user, etc.), and outputting first audio content using a second audio output mode includes outputting first audio content using a second set of sound sources located in a second set of locations in the physical environment, the second set of sound sources being different from the first set of sound sources. In some embodiments, the first set of sound sources and the second set of sound sources are housed in the same housing (e.g., the housing of the HMD, the housing of the same speaker or soundbar, etc.). In some embodiments, the first set of sound sources and the second set of sound sources are housed in different housings (for example, the first set of sound sources is surrounded by an HMD or earphones, and the second set of sound sources is surrounded by a set of external speakers positioned at various locations in the physical environment surrounding the user; the first set of sound sources is surrounded by a pair of speakers positioned in the physical environment around the user, and the second set of sound sources is surrounded by a set of three or more speakers positioned in the physical environment around the user, etc.). In some embodiments, the sound sources in the first set of sound sources and the second set of sound sources refer to elements of physical vibration that generate sound waves and propagate away from the location of the vibration elements. In some embodiments, the physical vibration characteristics of individual sound sources (e.g., wavefront shape, phase, amplitude, frequency, etc.) are controlled by a computer system according to the audio content output by the output device. In some embodiments, individual or individual subsets of sound sources in the first set of sound sources and / or the second set of sound sources have the same characteristics and different locations. In some embodiments, the individual or subset of sound sources within a first set and / or second set of sound sources have different characteristics and the same location.In some embodiments, the individual or subset of sound sources in a first set and / or a second set of sound sources have different characteristics and different locations. In some embodiments, the different characteristics of the individual sound sources or different subsets of sound sources in the first set and the second set of sound sources are individually controlled by the computer system based on the currently displayed portion of the first visual content and the corresponding audio content. In some embodiments, the sound sources in the first set of sound sources are not individually controlled (for example, the sound sources have the same phase, the same amplitude, the same wavefront shape, etc.). In some embodiments, sound sources within a second set of sound sources are individually controlled based on the spatial relationships between objects and virtual objects within the currently displayed portion of the first visual content (e.g., having different relative phases, different propagation directions, different amplitudes, different frequencies, etc.), and as a result, the sounds obtained at different locations in the physical environment are dynamically adjusted based on changes in the currently displayed portion of the first visual content (e.g., changes in the spatial relationships between objects within the currently displayed portion of the first visual content, different user interactions with different virtual objects or different parts of virtual objects within the currently displayed portion of the first visual content, different types of events occurring in the currently displayed portion of the first visual content, etc.).
[0225] Outputting first audio content using a first set of sound sources placed in a first set of locations within the physical environment, and outputting second audio content using a second set of sound sources different from the first set, placed in a second set of locations within the physical environment, provides the user with improved audio feedback (e.g., improved audio feedback regarding the current level of immersion). By providing improved feedback, the usability of the device is enhanced, and in addition, power consumption is reduced and battery life is improved by allowing the user to use the device more quickly and efficiently.
[0226] In some embodiments, the second set of sound sources includes the first set of sound sources and one or more additional sound sources not included in the first set of sound sources. In some embodiments, when the first visual content is displayed at a lower level of immersion and / or in a smaller spatial area (e.g., within a window or fixed frame), a smaller subset of sound sources (e.g., one or two sound sources, one or two sets of sound sources located in one or two locations, a sound source used to produce a single channel, or stereo sound, etc.) within the audio output device (one or more) associated with the computer system is used to output the first audio content. When the first visual content is displayed at a higher level of immersion and / or in a larger spatial extent (e.g., without a fixed window or fixed frame, spanning three-dimensional space surrounding the user, etc.), a larger subset or all of the available sound sources (e.g., three or more sound sources for producing surround sound and / or spatially positioned sounds, etc.) within the audio output device (one or more) associated with the computer system is used to output the first audio content. Outputting first audio content using a first set of sound sources, each placed in a first set of locations within the physical environment, and outputting second audio content using a second set of sound sources, each placed in a second set of locations within the physical environment, which includes the first set of sound sources and one or more additional sound sources not included in the first set, provides the user with improved audio feedback (e.g., improved audio feedback regarding the current level of immersion). By providing improved feedback, the usability of the device is enhanced, and in addition, power consumption is reduced and battery life is improved by allowing the user to use the device more quickly and efficiently.
[0227] In some embodiments, the second set of locations extends over a wider area than the first set of locations in the physical environment. In some embodiments, the first set of locations is positioned to the left and right of the user, or in front of the user. The second set of locations is positioned in three or more locations around the user (e.g., in front, left, right, behind, above, below, and / or optionally at other angles to the user's forward direction in three-dimensional space). Outputting first audio content using a first set of sound sources positioned in each of the first set of locations in the physical environment, and outputting second audio content using a second set of sound sources different from the first set of sound sources positioned in each of the second set of locations in the physical environment, which extends over a wider area than the first set of locations in the physical environment, provides the user with improved audio feedback (e.g., improved audio feedback regarding the current level of immersion). By providing improved feedback, the usability of the device is improved, and in addition, power consumption is reduced and the battery life of the device is improved by allowing the user to use the device more quickly and efficiently.
[0228] In some embodiments, outputting first audio content using a first audio output mode includes outputting the first audio content according to a pre-defined correspondence between the first audio content and first visual content (e.g., a temporal correspondence between an audio playback timeline and a video playback timeline, a pre-established content-based correspondence (e.g., a sound effect associated with an individual object, an alert associated with an individual user interface event, etc.)), wherein the pre-defined correspondence is independent of the spatial location of each virtual object in the currently displayed view of the first visual content (e.g., the spatial location of a virtual object in the currently displayed view of the first visual content is optional). Outputting the first audio content using a second audio output mode (which is modified according to the movement of virtual objects in the environment depicted in the first visual content and / or according to a changed viewpoint in the environment depicted by the three-dimensional environment, etc.) includes outputting the first audio content according to a pre-defined correspondence between the first audio content and the first visual content (e.g., a temporal correspondence between an audio playback timeline and a video playback timeline, a pre-established content-based correspondence (e.g., a sound effect associated with an individual object, an alert associated with an individual user interface event, etc.)) and according to the respective spatial locations of the virtual objects in the currently displayed view of the first visual content. For example, in some embodiments, when the first audio output mode is used to output the first audio content, the sound produced by the audio output device(s) is independent of the spatial relationships between the user's viewpoints corresponding to the currently displayed view of the first visual content. In some embodiments, when a first audio output mode is used to output first audio content, the sound produced by the audio output device(s) is independent of the spatial relationships between virtual objects in the currently displayed view of the first visual content.In some embodiments, when a first audio output mode is used to output first audio content, the sound produced by the audio output device(s) is independent of changes in the spatial relationships between virtual objects in the currently displayed view of the first visual content, caused by user input (e.g., when a virtual object that is a perceived producer of sound in the first visual content is moved by the user (e.g., in a user interface, game, virtual environment, etc.)). In some embodiments, when a first audio output mode is used to output first audio content, the sound produced by the audio output device(s) is headlocked to the user's head (e.g., when the user is wearing an HMD that includes an audio output device(s)), regardless of the user's viewpoint or spatial relationship to the virtual content presented in the computer-generated environment. In some embodiments, when a first audio output mode is used to output first audio content, the sound produced by the audio output device(s) is headlocked to the user's head (for example, when the user is wearing an HMD that includes the audio output device(s)) and is independent of the user's movements in the physical environment.
[0229] Outputting the first audio content according to a pre-configured correspondence between the first audio content and the first visual content, which is independent of the spatial location of each virtual object in the currently displayed view of the first visual content, and outputting the second audio content according to the pre-configured correspondence between the first audio content and the first visual content, and according to the spatial location of each virtual object in the currently displayed view of the first visual content, provides the user with improved audio feedback (e.g., improved audio feedback regarding the current level of immersion). By providing improved feedback, the usability of the device is improved, and in addition, power consumption is reduced and the battery life of the device is improved by allowing the user to use the device more quickly and efficiently.
[0230] In some embodiments, outputting the first audio content using a second audio output mode includes outputting a first portion of the first audio content corresponding to the currently displayed view of the first visual content having sound localization corresponding to a first spatial relationship, based on the determination that a first virtual object in the currently displayed view of the first visual content has a first spatial relationship with a viewpoint corresponding to the currently displayed view of the first visual content; and outputting a first portion of the first audio content corresponding to the currently displayed view of the first visual content with sound localization corresponding to a second spatial relationship, based on the determination that a first virtual object in the currently displayed view of the first visual content has a second spatial relationship with a viewpoint corresponding to the currently displayed view of the first visual content, wherein the first spatial relationship is different from the second spatial relationship, and the sound localization corresponding to the first spatial relationship is different from the sound localization corresponding to the second spatial relationship. For example, if the first visual content includes a singing bird and the corresponding first audio content includes the bird's song, the sound output in the second audio output mode is adjusted not only so that the volume changes based on the bird's perceived distance from the viewpoint of the currently displayed view, but also so that the perceived sound source changes depending on the bird's location relative to the viewpoint of the currently displayed view. In some embodiments, the perceived sound s...
Claims
1. It is a method, A computer system that communicates with a first display generation component and one or more input devices, Displaying a view of a three-dimensional environment via the first display generation component, The view of the three-dimensional environment simultaneously includes a first virtual content, a second virtual content different from the first virtual content, and a representation of a first part of the physical environment. The first portion of the physical environment includes a first physical surface, The first virtual content includes a first user interface object displayed at a position in the three-dimensional environment corresponding to the location of the first physical surface within the first portion of the physical environment, The second virtual content includes individual user interface objects displayed at separate positions in the three-dimensional environment that are different from the positions corresponding to the locations on the first physical surface, While displaying the view of the three-dimensional environment, the detection of a user's portion at a first location within the first portion of the physical environment, wherein the first location lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment. In response to detecting the portion of the user at a first location within the first portion of the physical environment, the display of the first portion of the first user interface object is stopped while maintaining the display of the second portion of the first user interface object, such that the representation of the portion of the user appears to be in the position where the first portion of the first user interface object was previously displayed, and at least the first portion of the representation of the portion of the user is visually obscured by the individual user interface object. While the view of the three-dimensional environment is being displayed, the movement of the user's portion from the first location to a second location within the first portion of the physical environment, wherein the second location lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment. In response to the detection of the movement of the user's portion from the first location to the second location, the display of the first portion of the first user interface object is restored and the display of the second portion of the first user interface object is stopped, such that the display of the first portion of the representation of the user's portion, which was visually obscured by the individual user interface object, is restored. Methods that include...
2. Detecting the user's portion at a first location within the first portion of the physical environment, and detecting a first input by the user's portion corresponding to a request to select the first user interface object, while maintaining the display of the second portion of the first user interface object without displaying the first portion of the first user interface object. In response to the detection of the first input by the user's part, the first operation corresponding to the first user interface object is performed. Detecting the user's portion at the second location within the first portion of the physical environment, and detecting a second input by the user's portion corresponding to the request to select the first user interface object, while maintaining the display of the first portion of the first user interface object without displaying the second portion of the first user interface object. The method according to claim 1, comprising: detecting the second input by the portion of the user; performing a second action corresponding to the first user interface object.
3. Detecting the user's portion at a first location within the first portion of the physical environment, and detecting a first input by the user's portion corresponding to a request to select the first portion of the first user interface object, while maintaining the display of the second portion of the first user interface object without displaying the first portion of the first user interface object. In response to the detection of the first input by the user's part, the first operation corresponding to the first part of the first user interface object is performed. Detecting the user's portion at the second location within the first portion of the physical environment, and detecting a second input by the user's portion corresponding to the request to select the second portion of the first user interface object, while maintaining the display of the first portion of the first user interface object without displaying the second portion of the first user interface object. The method according to claim 1, comprising: detecting the second input by the portion of the user; performing a second operation corresponding to the second portion of the first user interface object, wherein the second operation is different from the first operation.
4. The first virtual content includes a second user interface object displayed at the position in the three-dimensional environment corresponding to the location of the first physical surface within the first portion of the physical environment, and the method The method according to any one of claims 1 to 3, comprising, in response to detecting the portion of the user at a first location within the first portion of the physical environment, stopping the display of the first portion of the second user interface object while maintaining the display of the second portion of the second user interface object, so that the representation of the portion of the user appears to be in the position where the first portion of the second user interface object was previously displayed.
5. Detecting the user's portion at a first location within the first portion of the physical environment, and detecting a third input by the user's portion corresponding to a request to select the first user interface object, while maintaining the display of the second portion of the first user interface object and the second portion of the second user interface object without displaying the first portion of the first user interface object and the first portion of the second user interface object. The method according to claim 4, comprising: in response to detection of the third input by the portion of the user, performing a third operation corresponding to the first user interface object without performing a fourth operation corresponding to the second user interface object.
6. The method according to claim 4 or 5, comprising, in response to detection of the movement of the portion of the user from the first location to the second location, restoring the display of the first portion of the second user interface object and stopping the display of the second portion of the second user interface object so that the representation of the portion of the user appears to be in the position where the second portion of the second user interface object was previously displayed,
7. The method according to claim 4 or 5, comprising, in response to detection of the movement of the portion of the user from the first location to the second location, maintaining the display of the second portion of the second user interface object without restoring the display of the first portion of the second user interface object, such that the representation of the portion of the user appears to be in the position where the first portion of the second user interface object was previously displayed.
8. In response to the detection of the user's portion at a first location within the first portion of the physical environment, a simulated shadow of the user's portion is displayed at a first position in the view of the three-dimensional environment, offset from the position where the first portion of the first user interface object was previously displayed. The method according to any one of claims 1 to 7, comprising: detecting the movement of the portion of the user from the first location to the second location, and displaying the simulated shadow of the portion of the user at a second position in the view of the three-dimensional environment, offset from the position where the second portion of the first user interface object was previously displayed.
9. The method according to any one of claims 1 to 8, wherein the first user interface object is a virtual keyboard including at least a first key and a second key different from the first key, the first portion of the first user interface object corresponds to the first key, and the second portion of the first user interface object corresponds to the second key.
10. A computer system, Display generation components and One or more input devices, One or more processors, A memory storing one or more programs is provided, and the one or more programs are configured to be executed by the one or more processors, Displaying a view of a three-dimensional environment via the aforementioned display generation component, The view of the three-dimensional environment simultaneously includes a first virtual content, a second virtual content different from the first virtual content, and a representation of a first part of the physical environment. The first portion of the physical environment includes a first physical surface, The first virtual content includes a first user interface object displayed at a position in the three-dimensional environment corresponding to the location of the first physical surface within the first portion of the physical environment, The second virtual content includes individual user interface objects displayed at separate positions in the three-dimensional environment that are different from the positions corresponding to the locations on the first physical surface, While displaying the view of the three-dimensional environment, the detection of a user's portion at a first location within the first portion of the physical environment, wherein the first location lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment. In response to detecting the portion of the user at a first location within the first portion of the physical environment, the display of the first portion of the first user interface object is stopped while maintaining the display of the second portion of the first user interface object, such that the representation of the portion of the user appears to be in the position where the first portion of the first user interface object was previously displayed, and at least the first portion of the representation of the portion of the user is visually obscured by the individual user interface object. While the view of the three-dimensional environment is being displayed, the movement of the user's portion from the first location to a second location within the first portion of the physical environment, wherein the second location lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment. In response to the detection of the movement of the user's portion from the first location to the second location, the display of the first portion of the first user interface object is restored and the display of the second portion of the first user interface object is stopped, such that the display of the first portion of the representation of the user's portion, which was visually obscured by the individual user interface object, is restored. A computer system that includes instructions for performing a specific task.
11. The computer system according to claim 10, wherein one or more programs include instructions for performing the method described in any one of claims 2 to 9.
12. A computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by a computer system including a display generation component and one or more input devices, the computer system... Displaying a view of a three-dimensional environment via the aforementioned display generation component, The view of the three-dimensional environment simultaneously includes a first virtual content, a second virtual content different from the first virtual content, and a representation of a first part of the physical environment. The first portion of the physical environment includes a first physical surface, The first virtual content includes a first user interface object displayed at a position in the three-dimensional environment corresponding to the location of the first physical surface within the first portion of the physical environment, The second virtual content includes individual user interface objects displayed at separate positions in the three-dimensional environment that are different from the positions corresponding to the locations on the first physical surface, While displaying the view of the three-dimensional environment, the detection of a user's portion at a first location within the first portion of the physical environment, wherein the first location lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment. In response to detecting the portion of the user at a first location within the first portion of the physical environment, the display of the first portion of the first user interface object is stopped while maintaining the display of the second portion of the first user interface object, such that the representation of the portion of the user appears to be in the position where the first portion of the first user interface object was previously displayed, and at least the first portion of the representation of the portion of the user is visually obscured by the individual user interface object. While the view of the three-dimensional environment is being displayed, the movement of the user's portion from the first location to a second location within the first portion of the physical environment, wherein the second location lies between the first physical surface and the viewpoint corresponding to the view of the three-dimensional environment. In response to the detection of the movement of the user's portion from the first location to the second location, the display of the first portion of the first user interface object is restored and the display of the second portion of the first user interface object is stopped, such that the display of the first portion of the representation of the user's portion, which was visually obscured by the individual user interface object, is restored. A computer-readable storage medium containing instructions to perform an action.
13. The computer-readable storage medium according to claim 12, wherein one or more programs include instructions for performing the method according to any one of claims 2 to 9.
Citation Information
Patent Citations
Input device
JP1995005978A
Image processing method and image processing device
JP2006126936A
Information processing device and information processing method
JP2015041126A
Information processing apparatus and information processing method, and computer program
JP2016194744A
Hand-gesture-based interface utilizing augmented reality
US20160357263A1