Devices, methods, and graphical user interfaces for providing computer-generated experiences.

A dual-display system enhances virtual and augmented reality interaction by reducing the need for constant HMD use and providing real-time contextual information, improving efficiency and social interaction.

JP7911564B2Active Publication Date: 2026-08-26APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024171404
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-12
Filing Date
2024-09-30
Publication Date
2026-08-26
Estimated Expiration
2041-03-15

AI Technical Summary

Technical Problem

Existing methods and interfaces for interacting with virtual and augmented reality environments are cumbersome, inefficient, and limit social interaction, often requiring multiple inputs and causing errors, while head-mounted displays obstruct the user's view of the physical environment.

Method used

A computing system with dual display components, one facing the user for immersive content and another facing outward to provide real-time state information, reducing the need for constant HMD interaction and enhancing social cues.

Benefits of technology

Improves interaction efficiency, reduces errors, and facilitates social interaction by providing real-time contextual information to the user's surroundings, allowing seamless engagement with both virtual and physical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007911564000001
    Figure 0007911564000001
  • Figure 0007911564000002
    Figure 0007911564000002
  • Figure 0007911564000003
    Figure 0007911564000003
Patent Text Reader

Abstract

To provide a computer-generated experience to make interaction with a computing system more efficient and intuitive for a user.SOLUTION: A computing system displays, via a first display generation component, a first computer-generated environment and simultaneously displays, via a second display generation component, a visual representation of a portion of a user of the computing system who is in a position to view the first computer-generated environment via the first display generation component and one or more graphical elements that provide a visual indication of content in the first computer-generated environment. The computing system alters the visual representation of a portion of the user to represent a change in the appearance of the user over a discrete period and alters one or more graphical elements to represent a change in the first computer-generated environment over a discrete period.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] (Priority and Related Applications) This application is a continuation of U.S. Patent Application No. 17 / 200,676, filed on March 12, 2021, titled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR PROVIDING COMPUTER - GENERATED EXPERIENCES", which claims the priority of U.S. Provisional Application No. 62 / 990,408, filed on March 16, 2020.

Technical Field

[0002] This disclosure generally relates to a computing system having one or more display - generating components and one or more input devices for providing computer - generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via one or more displays.

Background Art

[0003] The development of computing systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or augment the representation of the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computing systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual three - dimensional objects, digital images, videos, text, icons, and control elements such as buttons and other graphics.

[0004] However, the methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems where manipulating virtual objects is complex and error-prone impair the user's cognitive burden and detract from the experience in the virtual / augmented reality environment. In addition, these methods are unnecessarily time-consuming and thereby waste energy. This latter consideration is particularly important in battery-powered devices. Furthermore, many systems that provide virtual reality and / or mixed reality experiences use head-mounted display devices that physically shield the user's face from their surroundings when the user engages with the virtual reality and mixed reality experience, hindering social interaction and information exchange with the outside world. [Overview of the project]

[0005] Therefore, there is a need for computing systems with improved methods and interfaces to provide users with computer-generated experiences that make interaction with the computing system more efficient and intuitive for the user. Furthermore, there is a need for computing systems with improved methods and interfaces to provide users with computer-generated experiences that facilitate better social interaction, etiquette, and information exchange with the surrounding environment while the user is engaged in various virtual and mixed reality experiences. Such methods and interfaces optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or types of user input by assisting the user in understanding the connection between the inputs provided and the device's response to those inputs, thereby generating a more efficient human-machine interface. Such methods and interfaces also improve the user experience, for example, by reducing errors, interruptions, and time delays caused by a lack of social cues and visual information on the user's and others' sides in the same physical environment when the user engages in virtual and / or mixed reality experiences provided by the computing system.

[0006] The above-mentioned defects and other problems relating to the user interface for a computing system having a display generation component and one or more input devices are mitigated or eliminated by the disclosed system. In some embodiments, the computing system is a desktop computer with one or more associated displays. In some embodiments, the computing system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computing system is a personal electronic device (e.g., a wearable electronic device such as a watch or head-mounted device). In some embodiments, the computing system has a touchpad. In some embodiments, the computing system has one or more cameras. In some embodiments, the computing system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computing system has one or more eye-tracking components. In some embodiments, the computing system has one or more hand-tracking components. In some embodiments, the computing system has one or more output devices in addition to one or more display generation components, the output devices include one or more tactile output generators and one or more audio output devices. In some embodiments, the computing system includes a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI or the user's body as captured by cameras and other motion sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally include non-temporary computer-readable storage media or other computer program products configured to be executed by one or more processors.

[0007] As disclosed herein, a computing system includes at least two display generation components, the first of which faces the user and provides the user with a three-dimensional computer-generated experience, and the second of which faces away from the user and provides state information related to the user (e.g., the user's eye movements) and / or state information related to the computer-generated experience currently being viewed by the user (e.g., metadata related to the content viewed by the user and the level of immersion associated with the content). The first and second display generation components are optionally two displays enclosed within the same housing of a head-mounted display device (HMD), facing inward toward the user wearing the HMD and outward toward the physical environment surrounding the user, respectively. The second display generation component optionally provides real-time state information including a visual representation of a portion of the user obscured behind the first display generation component, and metadata and / or associated immersion levels related to the content currently shown to the user via the first display generation component, thereby enabling another person(s) in the physical environment surrounding the user to see the visual information and metadata provided by the second display generation component and act accordingly, for example, engaging with the user when appropriate, rather than unnecessarily avoiding interaction with the user or inappropriately disturbing the user while the user is viewing computer-generated content via the first display generation component.In some embodiments, a user of a computing system optionally activates different modes of the computing system to suit the user's intended level of involvement and privacy needs when engaging with a computer-generated environment provided via a first display generation component, and the computing system provides state information related to the various modes to alert people in the surrounding physical environment of such intentions and needs of the user, thereby reducing the risk of unintended, unwanted, and / or unnecessary interruptions and interactions by people in the surrounding physical environment.

[0008] If a computing system includes at least two display generation components within the same housing as disclosed herein, a second (e.g., outward-facing) display generation component optionally displays contextual information indicating the availability of a computer-generated experience based on the current context. In response to detecting that a first (e.g., inward-facing) display generation component is positioned in front of the user (e.g., the user is wearing the HMD with the inward-facing display facing the user's eyes, or the user is holding the HMD with the inward-facing display in front of the user's eyes), the computing system provides the user with a computer-generated experience via the first display generation component. By automatically alerting the user to computer-generated experiences available via the outward-facing display based on the current context (e.g., while the user is positioned to view the outward-facing display (e.g., while the user is not wearing the HMD on their head, while the HMD is placed on a table, etc.)) and / or when the inward-facing display is positioned in front of the user's eyes (e.g., when the user is wearing the HMD on their head, or when the HMD is held with the inward display facing the user's face or eyes), the system reduces the number, complexity, and scope of inputs required for the user to discover what computer-generated experiences are available in various contexts and selectively view the desired computer-generated experience (e.g., without the need to constantly wear the HMD, and / or without the need to browse through selectable options to locate the desired CGR content item, and / or without the need to activate controls displayed while wearing the HMD to start the desired CGR experience).In some embodiments, depending on whether the first display generation component is actually worn by the user (e.g., secured to the user's head or body with straps, as opposed to being held in front of the user's eyes by the user's hands), the computing system optionally provides different computer-generated experiences corresponding to the wearing state of the first display generation component (e.g., displaying a preview of the available computer-generated experiences (e.g., a shortened, two-dimensional or three-dimensional, interactive, etc.) when the first display generation component is not actually worn by the user, and displaying the full version of the available computer-generated experiences when the first display generation component is worn by the user). By selectively displaying different versions or different computer-generated experiences of a computer-generated experience, depending not only on the position of the display generation component relative to the user (e.g., based on whether the position allows the user to view the CGR experience) but also on whether the display generation component is firmly worn by the user (e.g., based on whether the user's hands are free or necessary to hold the display generation component in its current position), the number of inputs required to trigger the intended result is reduced, the activation of the full computer-generated experience is avoided unnecessarily, thereby saving the user's time when they only wish to briefly preview the computer-generated experience, and saving battery power for the display generation component and computing system when powered by battery.

[0009] As disclosed herein, in some embodiments, the computing system includes a first display generating component and a second display generating component, which are located in the same housing or mounted on the same physical support structure. The first and second display generating components optionally have their respective display surfaces which are opaque and face opposite directions. The display generating components, together with the housing or support structure, can be quite bulky and may be cumbersome to attach to and detach from the user's head / body. The display generating components also together form a significant physical barrier between the user and others in the surrounding physical environment. By using an external display (e.g., a second display generation component) to show state information (e.g., title, progress, type, etc.) related to the metadata of the CGR content displayed on the internal display (e.g., a first display generation component), the immersion level associated with the displayed CGR content (e.g., full passthrough, mixed reality, virtual reality, etc.), and / or the visual characteristics of the displayed CGR content (e.g., changing color, brightness, etc.), the current display mode of the computing system (e.g., privacy mode, parental control mode, silent mode, etc.), and / or user characteristics (e.g., the appearance of the user's eyes, the user's identifier, etc.), the impact of the presence of physical barriers between the user and others in the surrounding environment is reduced without the user having to physically remove the display generation component, and unnecessary obstacles to desired social interaction and unnecessary interruptions to the user's engagement with the computer-generated experience are reduced. Furthermore, by using an external display to show contextual information and indications of contextually relevant computer-generated experiences, the user does not have to constantly pick up the HMD and place the internal display in front of their eyes to find out what CGR content is available. Users also do not need to be fully strapped into the HMD to preview available CGR experiences.The user only needs to fully wear the HMD when they wish to fully engage with the CGR experience (e.g., interacting with the CGR environment using aerial and micro-gestures). In this way, the number of times the user needs to position the HMD's internal display in front of their eyes and / or fully strap the HMD to their head is reduced without compromising the user's need to know what CGR experiences are available and / or hindering the user's ability to enjoy the desired CGR experience.

[0010] As disclosed herein, a computer-generated experience is provided via a display generation component of a computing system (e.g., a single display generation component of a device, an internal display of an HMD, etc.) in response to a user's physical interaction with a real-world physical object. Specifically, the computing system displays a visual indication that a computer-generated experience is available at a location in a three-dimensional environment displayed via the display generation component, the location of which corresponds to the location of a representation of the physical object in the three-dimensional environment. In response to detecting a physical interaction with a physical object in a first method that satisfies predefined criteria associated with the physical object, the computing system displays the physical object and, optionally, the computer-generated experience associated with the first method of physical interaction. For example, the computing system displays a pass-through view of the user's hand and the physical object before the predefined criteria are met by the user's operation of the physical object, and displays a computer-enhanced representation of the user's hand(s) manipulating the physical object after the predefined criteria are met. By automatically initiating a computer-generated experience in response to detecting a predefined physical interaction with a real-world physical object, the user's experience interacting with the physical object is improved, the interaction becomes more intuitive, and user errors when interacting with the physical object are reduced.

[0011] As disclosed herein, a computing system includes a display generation component (e.g., a single display generation component of a device, an internal display of an HMD) within a housing and provides a user interface (e.g., buttons, touch-sensitive surfaces) on the housing of the display generation component. When input is detected via the user interface, the computing system determines whether to perform an action associated with the input detected via the user interface on the housing of the display generation component, or to refrain from performing an action, depending on whether a pre-configured configuration of the user's hands touching the housing (e.g., both hands) is detected on the housing of the display generation component. By choosing to perform or refrain from performing an action in conjunction with the detection of input in conjunction with the hand configuration on the housing of the display generation component, the system automatically distinguishes between intentional user input and other touches on the housing of the display generation component for purposes other than providing input to trigger a specific action, helping to avoid unintended consequences, reducing user confusion, and making it faster and easier for the user to interact with the display generation component.

[0012] According to some embodiments, the method is performed in a computing system including a first display generation component, a second display generation component, and one or more input devices, and includes displaying a first computer-generated environment via the first display generation component, and while the first computer-generated environment is being displayed via the first display generation component, simultaneously displaying via the second display generation component a visual representation of a portion of a user of the computing system positioned to view the first computer-generated environment via the first display generation component, and one or more graphic elements that provide visual indications of the content within the first computer-generated environment, wherein the simultaneous display of the visual representation of the portion of the user and one or more graphic elements includes modifying the visual representation of the portion of the user to represent changes in the user's appearance over separate periods, and modifying one or more graphic elements that provide visual indications of the content within the first computer-generated environment to represent changes in the first computer-generated environment over separate periods.

[0013] According to some embodiments, the method is performed in a computing system comprising a first display generation component, a second display generation component, and one or more input devices, and includes: displaying a computer-generated environment via the first display generation component; displaying state information corresponding to the computing system via the second display generation component, which includes simultaneously displaying a visual representation of a portion of a user of the computing system positioned to view the computer-generated environment via the first display generation component, and one or more graphic elements that provide visual indications of the content within the computer-generated environment, while the computer-generated environment is being displayed via the first display generation component; detecting individual events; and, in response to the detection of individual events, changing the level of immersion of the computer-generated environment displayed via the first display generation component, and changing the appearance of the visual representation of a portion of the user of the computing system, which includes changing the state information displayed via the second display generation component.

[0014] According to some embodiments, the method is performed in a computing system including a first display generation component, a second display generation component, and one or more input devices, and comprises: displaying one or more user interface elements via the second display generation component; detecting that the first display generation component has been moved to a predetermined orientation relative to an individual part of the user while the one or more user interface elements are being displayed via the second display generation component; and, in response to the detection that the first display generation component has been moved to a predetermined orientation relative to an individual part of the user, displaying the first user interface element via the second display generation component when the first display generation component has been moved to a predetermined orientation relative to an individual part of the user. The system includes, in accordance with the determination that a computing system was in a certain state, displaying a first user interface via the first display generation component while the first display generation component is in a predetermined orientation relative to an individual part of the user, and, in accordance with the determination that a computing system was in a second state corresponding to displaying a second user interface element via the second display generation component instead of displaying the first user interface element via the second display generation component when the first display generation component is moved to a predetermined orientation relative to an individual part of the user, displaying a second user interface different from the first user interface via the first display generation component while the first display generation component is in a predetermined orientation relative to an individual part of the user.

[0015] According to some embodiments, the method is performed in a computing system including a first display generation component and one or more input devices, and includes detecting a first trigger event corresponding to the first display generation component being placed in a first default configuration for a user; providing a first computer-generated experience via the first display generation component in response to the detection of the first trigger event, according to a determination that the computing system including the first display generation component is being worn by the user while it is in a first default configuration for the user; and providing a second computer-generated experience, separate from the first computer-generated experience, via the first display generation component, according to a determination that the computing system including the first display generation component is not being worn by the user while it is in a first default configuration for the user.

[0016] According to some embodiments, the method is performed in a computing system including a first display generation component and one or more input devices, and includes: displaying a visual indication that a computer-generated experience corresponding to a physical object is available for display via the first display generation component while displaying a representation of a physical object at a location in a three-dimensional environment corresponding to the location of the physical object in a physical environment; detecting an interaction with a physical object in a physical environment while displaying the visual indication that a computer-generated experience is available for display via the first display generation component; displaying a computer-generated experience corresponding to a physical object via the first display generation component in response to the detection of an interaction with a physical object in a physical environment, according to a determination that the interaction with the physical object in a physical environment satisfies a first criterion corresponding to the physical object; and ceasing to display a computer-generated experience corresponding to a physical object according to a determination that the interaction with the physical object in a physical environment does not satisfy the first criterion.

[0017] According to some embodiments, the method is performed in a computing system including a housing, a first display generation component housed in the housing, and one or more input devices, and includes detecting a first hand on the housing housing the first display generation component, and, in conjunction with the detection of a second hand on the housing in response to the detection of a first hand on the housing housing the first display generation component, ceasing the execution of an action associated with the first hand in accordance with the determination that the first hand has been detected, and executing an action associated with the first hand in accordance with the determination that the first hand has been detected on the housing, without detecting another hand on the housing.

[0018] According to some embodiments, a computing system includes one or more display generating components (e.g., one or more displays, projectors, head-mounted displays, etc., enclosed in the same or different housings), one or more input devices (e.g., one or more cameras, a touch-sensing surface, one or more sensors optionally for detecting the intensity of contact with the touch-sensing surface), one or more tactile output generators optionally, one or more processors, and memory for storing one or more programs, the one or more programs being configured to be executed by the one or more processors, and the one or more programs including instructions to perform or cause to perform any of the operations described herein. According to some embodiments, a non-temporary computer-readable storage medium internally stores instructions, and when executed by a computing system having one or more display generating components, one or more input devices (e.g., one or more cameras, a touch-sensing surface, one or more sensors optionally for detecting the intensity of contact with the touch-sensing surface), and one or more tactile output generators optionally, causes the devices to perform or cause to perform any of the operations described herein. According to some embodiments, a graphical user interface on a computing system having one or more display generating components, one or more input devices (e.g., one or more cameras, a touch-sensing surface, or optionally one or more sensors for detecting the intensity of contact with the touch-sensing surface), optionally one or more tactile output generators, memory, and one or more processors for executing one or more programs stored in memory, includes one or more of the elements displayed in any of the methods described herein, and these elements are updated in response to input as described in any of the methods described herein.According to some embodiments, a computing system includes one or more display generation components, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, one or more sensors that optionally detect the intensity of contact with the touch-sensitive surface), one or more tactile output generators that optionally detect one or more tactile output generators, and means for performing or causing to perform any of the operations described herein. According to some embodiments, an information processing device for use in a computing system having one or more display generation components, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, one or more sensors that optionally detect the intensity of contact with the touch-sensitive surface), and one or more tactile output generators that optionally detect one or more tactile output generators, includes means for performing or causing to perform any of the operations described herein.

[0019] Accordingly, a computing system is provided having one or more display generation components with improved methods and interfaces for providing users with computer-generated experiences that make interaction with the computing system more efficient and intuitive for the user. Furthermore, a computing system is provided having improved methods and interfaces for providing users with computer-generated experiences that facilitate better social interaction, etiquette, and information exchange with the surrounding environment while the user is engaged in various virtual and mixed reality experiences. Such methods and interfaces optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or types of input from the user by assisting the user in understanding the connection between the inputs provided and the device's response to those inputs, thereby generating a more efficient human-machine interface. Such methods and interfaces also improve the user experience, for example, by reducing errors, interruptions, and time delays caused by the lack of social cues and visual information on the side of the user and others present in the same physical environment when the user engages in virtual and / or mixed reality experiences provided by the computing system.

[0020] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The functions and advantages described herein are not exhaustive, and many additional functions and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention. [Brief explanation of the drawing]

[0021] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, and like reference numerals refer to corresponding parts throughout the following figures.

[0022] [Figure 1] FIG. 1 is a block diagram showing an operating environment of a computing system for providing a CGR experience according to some embodiments.

[0023] [Figure 2] FIG. 2 is a block diagram showing a controller of a computing system configured to manage and adjust a user's CGR experience according to some embodiments.

[0024] [Figure 3] FIG. 3 is a block diagram showing a display generation component of a computing system configured to provide a visual component of a CGR experience to a user according to some embodiments.

[0025] [Figure 4] FIG. 4 is a block diagram showing a hand tracking unit of a computing system configured to capture a user's gesture input according to some embodiments.

[0026] [Figure 5] FIG. 5 is a block diagram showing an eye tracking unit of a computing system configured to capture a user's gaze input according to some embodiments.

[0027] [Figure 6] FIG. 6 is a flowchart showing a Glint-assisted gaze tracking pipeline according to some embodiments.

[0028] [Figure 7A]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays content to the user via the first display generation component while, according to some embodiments, the second display generation component displays dynamically updated state information associated with the user and / or content (e.g., representations corresponding to changes in the user's appearance behind the first display generation component, metadata of the content, changes in the level of immersion associated with content playback, etc.). [Figure 7B] The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays content to the user via the first display generation component while, according to some embodiments, the second display generation component displays dynamically updated state information associated with the user and / or content (e.g., representations corresponding to changes in the user's appearance behind the first display generation component, metadata of the content, changes in the level of immersion associated with content playback, etc.). [Figure 7C]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays content to the user via the first display generation component while, according to some embodiments, the second display generation component displays dynamically updated state information associated with the user and / or content (e.g., representations corresponding to changes in the user's appearance behind the first display generation component, metadata of the content, changes in the level of immersion associated with content playback, etc.). [Figure 7D] The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays content to the user via the first display generation component while, according to some embodiments, the second display generation component displays dynamically updated state information associated with the user and / or content (e.g., representations corresponding to changes in the user's appearance behind the first display generation component, metadata of the content, changes in the level of immersion associated with content playback, etc.). [Figure 7E]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays content to the user via the first display generation component while, according to some embodiments, the second display generation component displays dynamically updated state information associated with the user and / or content (e.g., representations corresponding to changes in the user's appearance behind the first display generation component, metadata of the content, changes in the level of immersion associated with content playback, etc.).

[0029] [Figure 7F]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays indications of different computer-generated experiences via the second display generation component based on contextual information associated with the computing system (e.g., the positions of the first and second display generation components, user identification information, current time, etc.) according to some embodiments, and triggers the display of different computer-generated experiences corresponding to the contextual information via the first display generation component in response to changes in spatial relationships (e.g., from not facing the user's eyes to facing the user's eyes, from being stationary on a table or in a bag to being raised to the user's eye level, etc.) and / or changes in the wearing state of the first display generation component on the user (e.g., from being supported by the user's hands to being supported by the user's head / nose / ears, from not being worn on the user's head / body to being worn on the user's head / body, etc.). In some embodiments, the computing system includes only a single display generation component and / or is configured in a pre-set configuration for the user, and does not display an indication of available computer-generated experiences before beginning to display different computer-generated experiences based on the wear state of the display generation component. [Figure 7G]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays indications of different computer-generated experiences via the second display generation component based on contextual information associated with the computing system (e.g., the positions of the first and second display generation components, user identification information, current time, etc.) according to some embodiments, and triggers the display of different computer-generated experiences corresponding to the contextual information via the first display generation component in response to changes in spatial relationships (e.g., from not facing the user's eyes to facing the user's eyes, from being stationary on a table or in a bag to being raised to the user's eye level, etc.) and / or changes in the wearing state of the first display generation component on the user (e.g., from being supported by the user's hands to being supported by the user's head / nose / ears, from not being worn on the user's head / body to being worn on the user's head / body, etc.). In some embodiments, the computing system includes only a single display generation component and / or is configured in a pre-set configuration for the user, and does not display an indication of available computer-generated experiences before beginning to display different computer-generated experiences based on the wear state of the display generation component. [Figure 7H]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays indications of different computer-generated experiences via the second display generation component based on contextual information associated with the computing system (e.g., the positions of the first and second display generation components, user identification information, current time, etc.) according to some embodiments, and triggers the display of different computer-generated experiences corresponding to the contextual information via the first display generation component in response to changes in spatial relationships (e.g., from not facing the user's eyes to facing the user's eyes, from being stationary on a table or in a bag to being raised to the user's eye level, etc.) and / or changes in the wearing state of the first display generation component on the user (e.g., from being supported by the user's hands to being supported by the user's head / nose / ears, from not being worn on the user's head / body to being worn on the user's head / body, etc.). In some embodiments, the computing system includes only a single display generation component and / or is configured in a pre-set configuration for the user, and does not display an indication of available computer-generated experiences before beginning to display different computer-generated experiences based on the wear state of the display generation component. [Figure 7I]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays indications of different computer-generated experiences via the second display generation component based on contextual information associated with the computing system (e.g., the positions of the first and second display generation components, user identification information, current time, etc.) according to some embodiments, and triggers the display of different computer-generated experiences corresponding to the contextual information via the first display generation component in response to changes in spatial relationships (e.g., from not facing the user's eyes to facing the user's eyes, from being stationary on a table or in a bag to being raised to the user's eye level, etc.) and / or changes in the wearing state of the first display generation component on the user (e.g., from being supported by the user's hands to being supported by the user's head / nose / ears, from not being worn on the user's head / body to being worn on the user's head / body, etc.). In some embodiments, the computing system includes only a single display generation component and / or is configured in a pre-set configuration for the user, and does not display an indication of available computer-generated experiences before beginning to display different computer-generated experiences based on the wear state of the display generation component. [Figure 7J]The present invention describes a computing system comprising a first display generation component and a second display generation component (e.g., separate displays facing different directions, displays enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). In some embodiments, the computing system displays indications of different computer-generated experiences via the second display generation component based on contextual information associated with the computing system (e.g., the positions of the first and second display generation components, user identification information, current time, etc.) according to some embodiments, and triggers the display of different computer-generated experiences corresponding to the contextual information via the first display generation component in response to changes in spatial relationships (e.g., from not facing the user's eyes to facing the user's eyes, from being stationary on a table or in a bag to being raised to the user's eye level, etc.) and / or changes in the wearing state of the first display generation component on the user (e.g., from being supported by the user's hands to being supported by the user's head / nose / ears, from not being worn on the user's head / body to being worn on the user's head / body, etc.). In some embodiments, the computing system includes only a single display generation component and / or is configured in a pre-set configuration for the user, and does not display an indication of available computer-generated experiences before beginning to display different computer-generated experiences based on the wear state of the display generation component.

[0030] [Figure 7K] This document demonstrates, in several embodiments, how to display an indication of the availability of a computer-generated experience at a location corresponding to a representation of a physical object in a mixed reality environment, and how to trigger the display of a computer-generated experience corresponding to a physical object in response to the detection of a pre-configured physical interaction with a real-world physical object. [Figure 7L] This document demonstrates, in several embodiments, how to display an indication of the availability of a computer-generated experience at a location corresponding to a representation of a physical object in a mixed reality environment, and how to trigger the display of a computer-generated experience corresponding to a physical object in response to the detection of a pre-configured physical interaction with a real-world physical object. [Figure 7M] This document demonstrates, in several embodiments, how to display an indication of the availability of a computer-generated experience at a location corresponding to a representation of a physical object in a mixed reality environment, and how to trigger the display of a computer-generated experience corresponding to a physical object in response to the detection of a pre-configured physical interaction with a real-world physical object.

[0031] [Figure 7N] This demonstrates that, in some embodiments, the display generation component can choose to perform or not perform an action depending on the input detected on the housing, depending on whether one hand or both hands are detected on the housing at the time the input is detected. [Figure 7O] This demonstrates that, in some embodiments, the display generation component can choose to perform or not perform an action depending on the input detected on the housing, depending on whether one hand or both hands are detected on the housing at the time the input is detected. [Figure 7P] This demonstrates that, in some embodiments, the display generation component can choose to perform or not perform an action depending on the input detected on the housing, depending on whether one hand or both hands are detected on the housing at the time the input is detected. [Figure 7Q] This demonstrates that, in some embodiments, the display generation component can choose to perform or not perform an action depending on the input detected on the housing, depending on whether one hand or both hands are detected on the housing at the time the input is detected.

[0032] [Figure 8]This is a flowchart of a method for displaying a computer-generated environment, state information associated with the computer-generated environment, and state information associated with a user viewing the computer-generated environment, according to several embodiments.

[0033] [Figure 9] This is a flowchart of a method for displaying a computer-generated environment, state information associated with the computer-generated environment, and state information associated with a user viewing the computer-generated environment, according to several embodiments.

[0034] [Figure 10] This is a flowchart illustrating a method for providing a computer-generated experience based on contextual information, according to several embodiments.

[0035] [Figure 11] This is a flowchart of a method for providing a computer-generated experience based on the wearing state of a display generation component, according to several embodiments.

[0036] [Figure 12] This is a flowchart of how to trigger the display of a computer-generated experience based on the detection of a pre-configured physical interaction with a real-world physical object, according to several embodiments.

[0037] [Figure 13] This is a flowchart illustrating how a display generation component performs actions in response to inputs on its housing, according to several embodiments. [Modes for carrying out the invention]

[0038] This disclosure relates to user interfaces that provide a computer-generated reality (CGR) experience to a user, in several embodiments.

[0039] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.

[0040] In some embodiments, the computing system includes a first display generation component and a second display generation component (e.g., separate displays, enclosed in the same housing but facing in different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). The first display generation component displays a computer-generated environment that provides a computer-generated experience to a user positioned to view content presented through the first display generation component (e.g., the user is facing the display surface of the display generation component (e.g., the surface of the physical environment illuminated by the projector, or the surface of the display emitting light that forms an image on the user's retina)). The first display generation component optionally provides a computer-generated experience with different levels of immersion, corresponding to different amounts of visual and audio information from the surrounding physical environment that are still perceptible through the first display generation component when the computer-generated experience is provided by the first display generation component. During normal operation (for example, when the user is wearing an HMD containing a first display-generating component and / or facing the display surface of the first display-generating component), the first display-generating component blocks the user's direct view of the surrounding physical environment and, at the same time, blocks the view of others of the user's face or eyes when the user is positioned to view content shown through the first display-generating component. In some embodiments, the first display-generating component is an internal display of the HMD that faces the user's eyes when the HMD is positioned on the user's head. Conventionally, when the user is positioned to view content shown through the display-generating component, the user has the option of seeing the physical environment or not seeing the physical environment by switching between displaying a computer-generated environment with different levels of appearance (for example, switching between full pass-through mode, mixed reality mode, or virtual reality mode).However, others in the surrounding environment facing the back of the display-generating component have little to no visual cues regarding the user's attention state, what content is displayed on the display-generating component, and / or whether the user can see the surrounding environment and the people within it. This imbalance of visual information (and optionally audio information) on the two faces of the display-generating component makes social interaction between the user and others in the surrounding environment unnatural and inefficient. Many considerations can benefit from a computing system that uses a second display-generating component to display an appropriate amount of visual information to people in the surrounding environment that conveys state information related to the user and / or the content displayed to the user via the first display-generating component. The display of state information by the second display-generating component is optionally displayed as long as the first display-generating component is in use, or optionally triggered in response to the detection of the presence of other people in the same physical environment and / or in response to the detection of an indication that the other person may wish to engage the user in a social conversation (e.g., by entering the same room, looking in the user's direction, waving to the user, etc.). In some embodiments, displaying state information on the second display generation component includes displaying a representation of a portion of the user (for example, a portion of the user that is obscured by the first display generation component when the user is in a position to view content displayed via the first display generation component) that is dynamically updated in accordance with changes in the user's appearance (for example, changes in a portion of the user that is obscured by the first display generation component). In some embodiments, displaying state information also includes displaying graphic elements that provide a visual indication of the content currently displayed via the first display generation component (for example, simultaneously with displaying a representation of a portion of the user).This method and system, which uses a second display generation component to display updated state information related to a user viewing content presented through the first display generation component, and metadata associated with the content state (e.g., title, progress, immersion level, display mode, etc.), allows others in the user's surrounding environment to gain useful insights into the user's current state while the user is engaged in a computer-generated experience, but without fully revealing the computer-generated experience to the surrounding environment. In some embodiments, representations of parts of the user obscured by the first display generation component (e.g., the user's eyes or face) and graphic elements indicating the state of content presented through the first display generation component are displayed on different display layers of the second display generation component, respectively, and are updated independently of each other. In some embodiments, the updates to representations of parts of the user and graphic elements indicating the state of content on different display layers of the second display generation component provide a more realistic view of the user's state behind a head-mounted display device housing both the first and second display generation components. The state information displayed on the second display generation component allows the user to remain socially connected to people in the surrounding environment while engaging in a computer-generated experience through the first display generation component. Dynamically updated state information on a second display generation component indicating the user's eye state and the state of the content presented to the user improves user engagement with computer-generated experiences when the user is in a public or semi-public environment by, for example, facilitating appropriate social interactions when such interactions are desired by the user; reducing unnecessary avoidance of social interactions by others in the surrounding environment due to a lack of visual cues about the user's permission to engage socially; informing others of appropriate times to interrupt the user's engagement with the computer-generated experience; and reducing unwelcome interruptions to the user's engagement experience due to a lack of visual cues about the user's desire to remain undisturbed.

[0041] As described above, many considerations can benefit from a computing system that uses a second display generation component to display an appropriate amount of visual information to other people in the surrounding environment that conveys state information related to the user and the content displayed to the user via the first display generation component. In some embodiments, the state information is displayed on the second display generation component as long as the first display generation component is in use. In some embodiments, the state information is displayed only in response to the detection of the presence of other people in the same physical environment and / or in response to the detection of some indication that others in the same physical environment may want the user to engage in social conversation (e.g., by entering the same room, looking in the user's direction, waving to the user, etc.). Displaying state information on the second display generation component optionally includes displaying a representation of a portion of the user (e.g., a portion of the user that is obscured by the first display generation component when the user is in a position to view content displayed via the first display generation component) and displaying graphic elements that provide visual indications of the content currently displayed via the first display generation component. Furthermore, in some embodiments, the representation of a portion of the user is updated in conjunction with changes in the level of immersion of the computer-generated experience displayed via the first display generation component. This method and system for updating state information, which includes using a second display generation component to display state information relating to and relating to the user viewing the content displayed via the first display generation component, and updating the appearance of the representation of a portion of the user in accordance with changes in the level of immersion associated with the content delivery, allows others in the user's surrounding environment to gain useful insights into the user's current state while the user is engaged in the computer-generated experience, without fully exposing the computer-generated experience to the surrounding environment.In some embodiments, updates to the representation of a portion of the user obscured by the first display generation component (e.g., the user's eyes or face), and updates to graphic elements indicating the state of the content displayed by the first display generation component, are shown on different display layers and updated independently of each other. By displaying the representation of the portion of the user and the graphic elements indicating the state of the content on different display layers, a more realistic view of the user's state behind a head-mounted display device housing both the first and second display generation components is provided. In some embodiments, the state information shown via the second display generation component (e.g., including graphic elements indicating the user's representation and the state of the content) optionally provides visual indications for many different usage modes of the computing system that address the different needs of the user and others in the same physical environment as the user. This allows the user to remain socially connected to people in their surrounding environment when engaging in a computer-generated experience. Dynamically updated state information on a second display generation component indicating the user's eye state and the state of the content presented to the user improves user engagement with computer-generated experiences when the user is in a public or semi-public environment by, for example, facilitating appropriate social interactions when such interactions are desired by the user; reducing unnecessary avoidance of social interactions by others in the surrounding environment due to a lack of visual cues about the user's permission to engage socially; informing others of appropriate times to interrupt the user's engagement with the computer-generated experience; and reducing unwelcome interruptions to the user's engagement experience due to a lack of visual cues about the user's desire to remain undisturbed.

[0042] In some embodiments, the computing system includes a first and second display generation component facing two different directions (e.g., separate displays, enclosed in the same housing but facing different directions (e.g., back-to-back facing opposite directions, or facing at different angles so that they cannot be viewed by the same user at the same time)). The first display generation component displays a computer-generated environment that provides the user with a computer-generated experience when the user comes to a position to view content presented through the first display generation component (e.g., facing the physical environment illuminated by a projector, or facing the display surface that emits light forming an image on the user's retina). The user may be in a position to view content presented on the second display generation component before the user positions the first display generation component to a position and orientation relative to the user viewing the content displayed on it (e.g., by moving the display generation component or the user themselves, or both). In an exemplary scenario, the first display-generating component is the internal display of the HMD that faces the user's eyes when the HMD is placed on the user's head, and the second display-generating component is the external display of the HMD that the user can view when the HMD is on a table or in the user's hand extended away from the user's face and not placed on the user's head or held close to the user's eyes. As disclosed herein, the computing system utilizes the second display-generating component to display an indication of the availability of different computer-generated experiences based on contextual information (e.g., location, time, user identification information, user permission level, etc.) and triggers the display of a selected computer-generated experience in response to detecting that the first display-generating component has been moved to a predetermined position and orientation relative to the user (e.g., the first display-generating component faces the user's eyes as a result of movement), enabling the user to view content shown via the first display-generating component.The displayed computer-generated experience is optionally selected based on the state of the second display-generating component at a point in time corresponding to the first display-generating component being moved to a predetermined position and orientation relative to the user. By indicating the availability of computer-generated experiences on the second display-generating component based on contextual information, and by automatically triggering the display of the selected computer-generated experience on the first display-generating component based on the state of the second display-generating component (and its contextual information) and changes in the orientation of the first display-generating component relative to the user, the system reduces the time and input required to achieve the desired result (e.g., retrieve information related to available experiences relevant to the current context and start the desired computer-generated experience), and reduces user errors and the time spent browsing and starting through available computer-generated experiences using a conventional user interface.

[0043] In some embodiments, the user can position and orient the first display-generating component relative to the user viewing the content displayed on it, in different ways, for example, in an impromptu or temporary way (e.g., held in front of the user at a distance, or held in the hand close to the user's eyes), or in a more formal and established way (e.g., secured to the user's head or face with a strap without being supported by the user's hands, or worn in another way). Depending on how the first display-generating component is positioned and oriented relative to the user, enabling the user to view the content displayed on the first display-generating component, the computing system selectively displays different computer-generated experiences (e.g., different versions of a computer-generated experience, different computer-generated experiences corresponding to different characteristics of the user or contextual properties, a preview of the experience versus the actual experience, etc.). In response to a trigger event corresponding to the placement of the first display generation component in a default configuration for the user, and according to how the first display generation component is held in its position and orientation (e.g., with or without the support of the user's hand, with or without the support of another mechanism other than the user's hand, etc.), the system selectively displays different computer-generated experiences (e.g., automatically starting the display of a computer-generated experience via the first display generation component without additional user input in the user interface provided by the first display generation component), thereby reducing the time and number of inputs required to achieve a desired result (e.g., starting a desired computer-generated experience), and reducing user errors and time spent browsing and starting through available computer-generated experiences using a conventional user interface.

[0044] In some embodiments, information (e.g., state information related to the user's eyes, the state of content displayed via the first display generation component, the display mode of the computing system, indications of available computer-generated experiences, etc.) is displayed on the second display generation component to help reduce the number of times the user needs to put on or take off the HMD containing both the first and second display generation components, and / or activate or deactivate computer-generated experiences, for example, to respond to others in the surrounding physical environment and / or to find a desired computer-generated experience. This helps save the user time, reduce power consumption, and reduce user errors when the user uses the display generation components, thereby improving the user experience. In some embodiments, a pre-configured method of physical manipulation of a real-world physical object is detected and used as a trigger to activate a computer-generated experience related to the physical object. In some embodiments, before activating a computer-generated experience related to a physical object, visual indications of available computer-generated experiences, and optionally, visual guides on how to activate a computer-generated experience (e.g., previews and animations), are displayed in the mixed reality environment at locations corresponding to the locations of the physical object's representation in the mixed reality environment. In addition to displaying visual indications regarding the availability of computer-generated experiences and / or visual guides regarding the physical actions required to trigger a computer-generated experience, the system enables users to achieve desired results (e.g., enter a desired computer-generated experience) more intuitively, quickly, and with less input by using pre-configured physical actions on physical objects to trigger the display of computer-generated experiences associated with those physical objects. This user interaction heuristic also helps reduce user errors when users interact with physical objects, thereby making the human-machine interface more efficient and saving power on battery-powered computing systems.

[0045] In some embodiments, the display generation component is housed in a housing that includes (or otherwise associated with) external sensors for detecting touch or hover input near or on various parts of the housing. Different types of touch and / or hover inputs (e.g., based on movement patterns (e.g., tap, swipe, etc.), duration (e.g., long, short, etc.), intensity (e.g., light, deep, etc.)) and different locations on or near the outside of the housing are used to trigger different actions associated with the display generation component or the computer-generated environment displayed by the display generation component. Interaction heuristics are used to determine whether an action should be performed depending on whether one hand or both hands are detected on the housing at the time the input is detected. By using the number of hands detected on the housing as an indicator of whether the user intends to provide input or is simply adjusting the position of the display generation component with the user's hands, it is possible to help reduce inadvertent or unintended actions of the display generation component, thereby making the human-machine interface more efficient and thus saving power on battery-powered computing systems.

[0046] Figures 1-6 illustrate exemplary computing systems for providing a CGR experience to a user. Figures 7A-7E show a computing system, in several embodiments, that displays content to a user via a first display generation component while displaying dynamically updated state information associated with the user and / or content via a second display generation component. Figures 7F-7J show a computing system, in several embodiments, that displays indications of different computer-generated experiences via a second display generation component based on contextual information, and optionally triggers the display of a different computer-generated experience corresponding to the contextual information via the first display generation component in response to detection of changes in spatial relationships to the user, and according to the wearing state of the first display generation component on the user. Figures 7K-7M show a computing system, in several embodiments, that displays indications of the availability of computer-generated experiences associated with physical objects in an augmented reality environment, and triggers the display of a computer-generated experience corresponding to the physical object in response to detection of a pre-configured physical interaction with the physical object. Figures 7N to 7Q show, in some embodiments, the display generation component selects whether to perform or not perform an action associated with the input detected on the housing, based on a determination of whether one hand or both hands are detected on the housing when the input is detected on the housing. Figure 8 is a flowchart of a method for displaying a computer-generated environment and state information in some embodiments. Figure 9 is a flowchart of a method for displaying a computer-generated environment and state information in some embodiments. Figure 10 is a flowchart of a method for providing a computer-generated experience based on contextual information in some embodiments. Figure 11 is a flowchart of a method for providing a computer-generated experience based on the wearing state of the display generation component in some embodiments. Figure 12 is a flowchart of a method for triggering the display of a computer-generated experience based on physical interaction with a physical object in some embodiments.Figure 13 is a flowchart illustrating how, in several embodiments, an operation is performed in response to input detected on the housing of a display generation component. The user interfaces in Figures 7A to 7Q are used to illustrate each of the processes in Figures 8 to 13.

[0047] In some embodiments, as shown in Figure 1, the CGR experience is provided to the user via an operating environment 100 that includes a computing system 101. The computing system 101 includes a controller 110 (e.g., a processor for a portable electronic device or remote server), one or more display generation components 120 (e.g., one or more head-mounted devices (HMDs) enclosed in the same housing and facing different directions, or enclosed in separate housings, HMDs having an internal display and an external display, one or more displays, one or more projectors, one or more touchscreens, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with a display generation component 120 (for example, in a head-mounted device (e.g., on the housing of an HMD or a display facing outwards from the HMD) or a handheld device).

[0048] When describing a CGR experience, various terms are used to refer individually to several related but distinct environments that the user can perceive and / or interact with (for example, using inputs detected by the computing system 101 that generates the CGR experience, causing the computing system that generates the CGR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computing system 101). The following is a subset of these terms.

[0049] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.

[0050] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In CGR, a subset of a person's bodily movements or their representations are tracked, and in response, one or more properties of one or more virtual objects simulated within the CGR environment are adjusted to behave according to at least one law of physics. For example, a CGR system may detect a person's head rotation and, in response, adjust the graphic content and sound field presented to the person in a similar manner to how such views and sounds would change in a physical environment. Depending on the circumstances (e.g., for reasons of accessibility), adjustments to the properties(s) of virtual objects(s) in the CGR environment may be made in response to representations of bodily movements (e.g., voice commands). A person may perceive and / or interact with CGR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may perceive and / or interact with an audio object that creates a 3D or spatially expansive audio environment, providing the perception of a point source in 3D space. In another example, an audio object may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, a person may perceive and / or interact with only audio objects.

[0051] Examples of CGR include virtual reality and mixed reality.

[0052] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their bodily movements within the computer-generated environment.

[0053] Mixed Reality: A mixed reality (MR) environment, in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input, refers to a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects). On a virtual continuum, a mixed reality environment is any place between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track the position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.

[0054] Examples of mixed reality include augmented reality and augmented virtual reality.

[0055] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through images or videos of the physical environment. As used herein, videos of the physical environment shown on an opaque display are referred to as “pass-through videos,” and it means that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, thereby allowing a person to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the image sensor. As another example, the representation of the physical environment may be transformed by graphically altering (e.g., enlarging) a portion of it, thereby making the altered portion a modified version that represents the original captured image but is not photorealistic. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.

[0056] Augmented Virtuality (AV): An augmented virtuality (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, but people with faces might be realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.

[0057] Hardware: There are many different types of electronic systems that enable people to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to the human eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination thereof. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the human retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces. In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience.In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., a physical setup / environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is communicably coupled to a display generation component 120 (e.g., one or more HMDs, displays, projectors, touchscreens, etc., enclosed in the same or different housings) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE802.11x, IEEE802.16x, IEEE802.3x, etc.). In another embodiment, the controller 110 is housed in a housing (e.g., a physical housing) containing one or more of the following: display generation components 120 (e.g., a portable electronic device including a display and one or more processors), one or more input devices 125, one or more output devices 155, one or more sensors 190, and / or one or more peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.

[0058] In some embodiments, at least one of the display generation components 120 is configured to provide the user with a CGR experience (e.g., at least the visual components of the CGR experience). In some embodiments, the display generation components 120 include a preferred combination of software, firmware, and / or hardware. One embodiment of the display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation components 120.

[0059] According to some embodiments, at least one of the display generation components 120 provides the user with a CGR experience while the user is virtually and / or physically present in scene 105.

[0060] In some embodiments, the display generation component(s) is worn on a part of the user's body (e.g., the user's head, the user's hand, etc.). Thus, at least one of the display generation component(s) 120 includes one or more CGR displays provided for displaying CGR content. For example, at least one of the display generation component(s) 120 surrounds the user's field of view. In some embodiments, at least one of the display generation component(s) 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds a device having a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, at least one of the display generation components 120 (one or more) is a CGR chamber, housing, or room configured to present CGR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with CGR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the CGR content response is displayed via the HMD.Similarly, a user interface that demonstrates interaction with CGR content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes (single or multiple), head, or hands)) can be implemented in the same way as an HMD triggered by the movement of the HMD relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes (single or multiple), head, or hands)).

[0061] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features for the sake of simplification are not shown so as not to obscure more suitable embodiments of the exemplary embodiments disclosed herein.

[0062] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0063] In some embodiments, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0064] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and CGR experience module 240.

[0065] The operating system 230 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). For this purpose, in various embodiments, the CGR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0066] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least one of the display generation components 120 1 in Figure 1, and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0067] In some embodiments, the tracking unit 242 is configured to map scene 105 to scene 105 in Figure 1 and to track the position / location of at least one of the display generation components 120 relative to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the tracking unit 242 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position of one or more parts of the user's hand relative to at least one of the display generation components 120 relative to scene 105 in Figure 1, and / or relative to a coordinate system defined for the user's hand, and / or relative to one or more parts of the user's hand. The hand tracking unit 244 is described in more detail below with reference to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to CGR content displayed via at least one of the display generation components 120. The eye-tracking unit 243 is described in more detail below with reference to Figure 5.

[0068] In some embodiments, the adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by at least one of the display generation components 120, and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0069] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least one of the display generation components 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0070] While the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., a controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.

[0071] Furthermore, Figure 2 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary from embodiment to embodiment and in some embodiments will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0072] Figure 3 is a block diagram of at least one example of one or more display generation components 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting embodiments, a computing system (e.g., HMD) including one or more display generation components 120 also includes, within the same housing, one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0073] In some embodiments, one or more communication buses 304 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).

[0074] In some embodiments, one or more CGR displays 312 are configured to provide the user with a CGR experience and optionally state information related to the CGR experience. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, an HMD includes a single CGR display. In another embodiment, an HMD includes a CGR display for each of the user's eyes. In some embodiments, one or more CGR displays 312 can present MR or VR content. In some embodiments, the HMD includes one or more CGR displays facing the user's eyes when the HMD is positioned on the user's head, and one or more CGR displays facing away from the user's eyes (for example, towards the external environment). In some embodiments, the computing system is a CGR room or CGR enclosure, which includes an internal CGR display that provides CGR content to a user inside the CGR room or enclosure, and optionally includes one or more external peripheral displays that display state information related to the CGR content and the state of the user inside.

[0075] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user is viewing when no display generation component(s) 120 are present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., with complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.

[0076] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and CGR presentation module 340.

[0077] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to the user via one or more CGR displays 312. To this end, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, a data transmission unit 348, and optionally other operational units for displaying status information related to the user and the CGR content.

[0078] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0079] In some embodiments, the CGR presentation unit 344 is configured to present CGR content and associated state information via one or more CGR displays 312. For this purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0080] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (for example, a 3D map of a mixed reality scene or a map of a physical environment on which computer-generated objects can be placed to generate computer-generated reality) based on media content data. For this purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0081] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0082] Although the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.

[0083] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary from embodiment to embodiment and in some embodiments will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0084] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by a hand tracking unit 244 (Figure 2) to track the position and / or movement of one or more parts of the user's hand relative to the scene 105 of Figure 1 (e.g., relative to a part of the physical environment surrounding the user, to at least one of the display generation components 120, or to a part of the user (e.g., the user's face, eyes, or head), and / or to a coordinate system defined for the user's hand). In some embodiments, the hand tracking device 140 is part of at least one of the display generation components 120 (e.g., embedded in or mounted in the same housing as the display generation components 120 (e.g., in a head-mounted device)). In some embodiments, the hand tracking device 140 is separate from the display generation components 120 (e.g., located in a separate housing or mounted in a separate physical support structure).

[0085] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.

[0086] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an Application Program Interface (API), which in turn drives one or more display generation components 120. For example, a user can interact with software running on the controller 110 by moving their hand 408 and changing the hand's posture.

[0087] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the hand tracking device 440 may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.

[0088] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D positions of the user's wrist and fingertips.

[0089] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The posture estimation function described herein may be interleaved with the motion tracking function, so that patch-based posture estimation is performed only once every two (or more) frames, while tracking is used to detect changes in posture that occur over the remaining frames. Posture, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 in response to the posture and / or gesture information, or perform other functions.

[0090] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be executed on dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 440, but some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the hand tracking device 402, or in other ways. In some embodiments, at least some of these processing functions may be executed by a suitable processor integrated with a display generation component (one or more) 120 (e.g., in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.

[0091] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values ​​to identify and segment image components (i.e., adjacent pixel groups) that have features of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.

[0092] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the position and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.

[0093] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 244 (Figure 2) to track the position and movement of the user's gaze toward the scene 105 or toward the CGR content displayed via at least one of the display generation components 120. In some embodiments, the eye-tracking device 130 is integrated with at least one of the display generation components 120. For example, in some embodiments, if the display generation components 120 are part of a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating CGR content for user viewing and a component for tracking the user's gaze toward the CGR content. In some embodiments, the eye-tracking device 130 is separate from the display generation components 120. For example, if the display generation component(s) is provided by a handheld device or CGR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or CGR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with at least one of the head-mounted display generation components(s) or at least one of the non-head-mounted display generation components(s). In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally part of a non-head-mounted display generation component.

[0094] In some embodiments, at least one of the display generation components (one or more) 120 uses a display mechanism (e.g., display panels near the left and right eyes) that displays frames containing left and right images in front of the user's eyes, and thus provides the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, at least one of the display generation components (one or more) 120 may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, at least one of the display generation components (one or more) 120 projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that an individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0095] As shown in Figure 5, in some embodiments, the eye-tracking device 130 includes at least one eye-tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visual light to pass through. The eye-tracking device 130 optionally captures images of the user's eye (e.g., as a video stream captured at 60 to 120 frames per second (fps)), analyzes the images to generate eye-tracking information, and communicates the eye-tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a corresponding eye-tracking camera and light source.

[0096] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR device to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil position, central visual position, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a glint-assisted method to determine the user's current visual axis and viewpoint relative to the display.

[0097] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece (one or more) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes (one or more) 592. The eye-tracking camera 540 may be directed towards a mirror 550 located between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, an internal display of a head-mounted device, or a display or projector of a handheld device), which transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or it may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the user's eye(s) 592 (e.g., as shown at the bottom of Figure 5).

[0098] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses eye-tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the eye-tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the eye-tracking input 542 is optionally used to determine the direction the user is currently looking.

[0099] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the CGR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can utilize the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.

[0100] In some embodiments, the eye-tracking device is part of a head-mounted device mounted on a wearable housing, which includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., a light source 530 (e.g., IR or NIR LEDs)). The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circular pattern around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.

[0101] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the position and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wide field of view (FOV) and a camera 540 with a narrow FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0102] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality (including, for example, virtual reality and / or mixed reality) applications to provide users with computer-generated reality (including, for example, virtual reality, augmented reality and / or augmented virtual reality) experiences.

[0103] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in a tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in a tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.

[0104] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which is initiated at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0105] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.

[0106] At 640, if the process proceeds from element 410, the current frame is analyzed and the pupil and glint are tracked, based in part on prior information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to confirm that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.

[0107] Figure 6 is intended to serve as an example of an eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in place of, or in combination with, the glint-assisted vision eye-tracking technology described herein in the computing system 101 to provide the user with a CGR experience in various embodiments.

[0108] This disclosure describes various input methods for interaction with a computing system. Where one example is provided using one input device or method, and another example is provided using a different input device or method, it should be understood that each example is compatible with the input device or method described in the other example and can be used optionally. Similarly, various output methods for interaction with a computing system are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, it should be understood that each example is compatible with the output device or method described in the other example and can be used optionally. Similarly, various methods for interaction with a virtual or mixed reality environment via a computing system are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with the method described in the other example and can be used optionally. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each embodiment. User interface and related processes

[0109] Here, we focus on embodiments of user interfaces ("UI") and related processes that can be executed on a computing system such as a portable multifunction device or head-mounted device, which comprises one or more display generation components, one or more input devices, and (optionally) one or a camera.

[0110] Figures 7A to 7E show a computing system (such as computing system 101 in Figure 1 or computing system 140 in Figure 4) according to several embodiments, which includes at least a first display generation component (e.g., display 7100) and a second display generation component (e.g., display 7102), and which displays computer-generated content to the user via the first display generation component (e.g., display 7100) while displaying dynamically updated state information associated with the user and / or content via the second display generation component (e.g., display 7102). Figures 7A to 7E are used to illustrate a process described later, including the process in Figures 8 to 13.

[0111] As shown in the left portion of Figure 7A, the first display generation component (e.g., display 7100) is located at position A7000-a and displays CGR content (e.g., a three-dimensional movie, virtual reality game, video, a three-dimensional environment including user interface objects, etc.). The first user 7202 is also located at position A7000-a.

[0112] As shown in the right-hand portion of Figure 7A, a second display generation component (e.g., display 7102) is located at position B7000-b and displays state information corresponding to the CGR content presented via the first user 7202 and / or the first display generation component (e.g., display 7100). In the exemplary scenario shown in Figure 7A, the second user 7204 is also located at position B7000-b.

[0113] As shown in Figure 7A, the spatial relationship between the first display generation component (e.g., display 7100) and the first user 7202 is such that the first user 7202 is positioned to view CGR content presented via the first display generation component. For example, the first user 7202 is facing the display surface of the first display generation component. In some embodiments, the first display generation component is the internal display of the HMD, and the spatial relationship represented by the simultaneous presence of the display 7100 and the first user 7202 at the same position A7000-a corresponds to the first user wearing or holding the HMD with the internal display of the HMD facing the user's eyes. In some embodiments, the first user is positioned to view CGR content presented via the first display generation component when the first user is facing a portion of the physical environment illuminated by the projection system of the first display generation component. For example, virtual content is projected onto a portion of the physical environment, and the virtual content and the portion of the physical environment are visible to the user through a camera view of the portion of the physical environment or through a transparent portion of the first display-generating component when the user is facing the display surface of the first display-generating component. In some embodiments, the first display-generating component emits light that forms an image on the user's retina when the user is facing the display surface of the first display-generating component. For example, virtual content is displayed by an LCD or LED display, superimposed on or replacing a portion of the view of the physical environment displayed by the LCD or LED display, and a user facing the display surface of the LCD or LED display can see the virtual content together with the view of the portion of the physical environment. In some embodiments, the first display-generating component displays a camera view of the physical environment in front of the first user, or includes a transparent or translucent portion of the physical environment in front of the first user that is visible to the first user.In some embodiments, a portion of the physical environment made visible to the first user via the first display generation component is a portion of the physical environment corresponding to the display surface of the second display generation component 7102 (e.g., the display surface of the second display generation component and optionally position B7000-b including the second user 7204). In some embodiments, the display surface of the second display generation component is a surface of the second display generation component facing away from the first user when the first user is in a position to view content shown by the first display generation component (e.g., when the first user is facing the display surface of the first display generation component), and is a light-emitting surface that forms an image visible to others facing a predetermined portion of the first user (e.g., the second user 7204 or others facing the face or eyes of the first user in the physical environment).

[0114] As shown in Figure 7A, the spatial relationship between the second display generation component (e.g., display 7102) and the second user 7204 is such that the second user 7204 is positioned to view state information presented by the second display generation component. For example, the second user 7204 is in front of and / or facing the display surface of the second display generation component. In some embodiments, the second display generation component is an external display of the HMD, including an internal display (e.g., represented by display 7100) that presents CGR content to the first user 7202. In such embodiments, the spatial relationship represented by the simultaneous presence of display 7102 and the second user 7204 at the same position B7000-b corresponds to the second user being in a portion of the physical environment that the external display of the HMD faces (e.g., the physical environment also accepts the first display generation component and the first user 7202). In some embodiments, the first display generation component displays a camera view of the physical environment in front of the first user, or includes a transparent or translucent pass-through portion that allows the first user to see a portion of the physical environment in front of the first user, and the portion of the physical environment included in the camera view or pass-through portion is also a portion of the physical environment in front of the display surface of the second display generation component. In some embodiments, the second display generation component is positioned back-to-back with the first display generation component such that a portion of the physical environment in front of the display surface of the second display generation component 7102 (e.g., the display surface of the second display generation component and optionally position B7000-b including the second user 7204) is also in front of the first user and is within the first user's field of view if the first and second display generation components do not obstruct the first user's face.

[0115] As stated above and repeatedly herein, Figure 7A (and Figures 7B-7J) shows the first display generation component (e.g., display 7100) and the second display generation component (e.g., display 7102) located in two separate and isolated parts of the physical environment, but it should be understood that the first and second display generation components are two display generation components that are optionally housed in the same housing (e.g., the housing of a single HMD) or mounted on the same support structure (e.g., back-to-back or mounted on two faces of a single wall or surface) and facing different directions (e.g., substantially opposite). Therefore, location A7000-a represents a first portion of the physical environment in which content presented via the first display generation component (e.g., CGR content) can be seen by a first user (e.g., first user 7202) facing the display surface of the first display generation component, and content presented via the second display generation component (e.g., state information) cannot be seen by the first user (e.g., first user 7202). Location B7000-b represents a second portion of the same physical environment in which content presented via the first display generation component (e.g., CGR content) cannot be seen by another user (e.g., second user 7204) facing the display surface of the second display generation component, and content presented via the second display generation component (e.g., state information) can be seen by that other user (e.g., second user 7204).In the disclosures presented herein, the first and second display generation components are controlled by the same computing system (e.g., an HMD, a portable electronic device housed separately from the display generation components, a portable electronic device having two displays facing in different directions, a remote server computer, etc.), and the user of the computing system generally refers, unless otherwise specified, to a person who has control of at least the first display generation component and positions it or himself in a position that enables him to view the CGR content shown through the first display generation component.

[0116] As shown in Figure 7A, the computing system controlling the first and second display generation components communicates with the first image sensor (e.g., camera 7104) and the second image sensor (e.g., camera 7106). The first image sensor faces the display surface of the first display generation component (e.g., display 7100) and is configured to capture an image of a portion of the physical environment (e.g., location A7000-a) that includes at least a portion of the first user (e.g., the face and / or eyes of the first user 7202) and does not include the second user 7204 (or, if the first display generation component is an internal display of an HMD worn by the first user, any other user). The second image sensor is configured to capture an image of a portion of the physical environment (e.g., location B7000-b) that does not include a portion of the first user (e.g., the face or eyes of the first user 7202), but includes at least a portion of the second user (e.g., a portion of the second user 7204 within the field of view of the first user 7202 provided by the first display generation component). As described above, in some embodiments, the portion of the physical environment captured by the second image sensor 7106 includes a portion of the physical environment within the field of view of the first user, provided that the eyes of the first user are not physically obstructed by the presence of the second display generation component (and optionally, the presence of the first display generation component). Similarly, in some embodiments, the portion of the physical environment captured by the first image sensor 7104 includes a portion of the user (e.g., the user's face or eyes) that is physically obstructed by the presence of the first display generation component (and optionally, the presence of the second display generation component). In some embodiments, the computing system also communicates with a first image sensor, a second image sensor, and / or other image sensors to receive images of the hands and wrists of the first and / or second users in order to identify gesture inputs provided by the first and / or second users. In some embodiments, the first image sensor 7104 is also used to capture gaze input provided by the first user.In some embodiments, the first and second image sensors optionally function as image sensors for capturing gesture input from a first user and / or a second user.

[0117] In some embodiments, the computing system controls one or more audio output devices that optionally provide audio output (e.g., sound of CGR content) to a first user located at location A7000-a, and optionally provide audio output (e.g., state indication sound or alarm, sound of CGR content, etc.) to a second user located at location B7000-b. In some embodiments, the computing system optionally partially or completely shields location A and the first user from sound propagating from location B (e.g., via one or more active or passive noise suppression or cancellation components), and optionally partially or completely shields location B and the second user from sound propagating from location A. In some embodiments, the amount of active sound insulation or sound pass-through is determined by a computing system based on the current level of immersion associated with the CGR content shown via the first display generation component (e.g., no sound insulation when in pass-through mode, partial sound insulation when in mixed reality mode, full sound insulation when in virtual reality mode, etc.) and optionally based on whether there is another user present at location B (e.g., no sound insulation when no one is present at location B, sound insulation when people are present at location B or the noise level exceeds a threshold level, etc.).

[0118] In some embodiments, as shown in Figure 7A, the computing system displays CGR content 7002 (e.g., shown as 7002-a in Figure 7A) via a first display generation component (e.g., display 7100, or the internal display of the HMD) while the first user 7202 is in a position to view CGR content (e.g., the first user 7202 is positioned juxtaposed with the display surface of the first display generation component at position A and at least partially facing it, the first user is wearing an HMD on their head, the first user is holding the HMD with the internal display in front of the user's eyes, etc.). At the moment shown in Figure 7A, the computing system is displaying a movie X (e.g., a three-dimensional movie, a two-dimensional movie, an interactive computer-generated experience, etc.). The movie is displayed in mixed reality mode, where the movie content is seen simultaneously with a representation of the physical environment (e.g., a representation of position B (e.g., a portion of the physical environment in front of the first user that is obscured by the presence of the first display generation component) via the first display generation component. In some embodiments, this mixed reality mode corresponds to an intermediate immersion level associated with CGR content presented via the first display generation component. In some embodiments, the intermediate immersion level also corresponds to partial occlusion or partial pass-through of sound propagating from the physical environment (e.g., position B (e.g., a portion of the physical environment surrounding the first user)). In this embodiment, the representation of the physical environment includes a representation 7010 (e.g., shown as 7010-a in Figure 7A) of a second user 7204 located at position B7000-b (e.g., in front of the second display generation component 7102 and also in front of the back of the first display generation component 7100). In some embodiments, the representation of the physical environment includes a camera view of a portion of the physical environment that is within the field of view of the first user if the user's eyes were not obstructed by the presence of the first and second display generation components (e.g., if the first user was not wearing an HMD or was not holding the HMD in front of the user's eyes).In mixed reality mode, CGR content 7002 (e.g., movie X, a three-dimensional augmented reality environment, user interface, virtual object, etc.) is displayed so as to superimpose or replace at least a portion, but not all, of the representation of the physical environment. In some embodiments, the first display generation component includes a transparent portion into which a portion of the physical environment is visible to the first user. In some embodiments, in mixed reality mode, the CGR content 7002 (e.g., movie X, a three-dimensional augmented reality environment, user interface, virtual object, etc.) is projected onto a physical surface or empty space within the physical environment and is visible through the transparent portion along with the physical environment, is visible through the transparent portion of the first display generation component, or is visible through a camera view of the physical environment provided by the first display generation component. In some embodiments, the CGR content 7002 is displayed so as to superimpose on a portion of the display, blocking the view of at least a portion, but not all, of the physical environment visible through the transparent or translucent portion of the first display generation component. In some embodiments, the first display generation component 7100 does not provide a view of the physical environment, but rather provides a complete virtual environment (e.g., without a camera view or transparent pass-through portion) augmented with a real-time visual representation (one or more) of the physical environment (e.g., a stylized representation or segmented camera image) as currently captured by one or more sensors (e.g., a camera, motion sensor, other attitude sensor, etc.). In mixed reality modes (e.g., augmented reality based on a camera view or transparent display, or augmented virtual based on a virtualized representation of the physical environment), the first user is not fully immersed in the computer-generated environment, but still receives sensory information (e.g., visual and audio) that directly corresponds to the physical environment surrounding the first user and the first display generation component.

[0119] As shown in Figure 7A, the computing system displays CGR content 7002-a (e.g., movie X, a three-dimensional augmented reality environment, user interface, virtual objects, etc.) in mixed reality mode via a first display generation component 7100, while simultaneously the computing system displays state information related to the first user and the CGR content via a second display generation component (e.g., display 7102, or an external display of the HMD). As shown in the right-hand portion of Figure 7A, the second display generation component (e.g., display 7102, or an external display of the HMD) displays one or more graphic elements representing the state of the CGR content 7002 displayed via the first display generation component (e.g., display 7100 or the internal display of the HMD), as well as a representation 7006 (e.g., shown as 7006-a in Figure 7A) of at least a portion of the first user 7202 in front of the display surface of the first display generation component (e.g., display 7100 or the internal display of the HMD). In this embodiment, one or more graphic elements representing the state of the CGR content displayed via the first display generation component optionally include an identifier for the CGR content (e.g., the title of movie X), a progress bar 7004 indicating the current progress of the CGR content (e.g., shown as 7004-a in Figure 7A), and a visual representation of the CGR content 7008 (e.g., shown as 7008-a in Figure 7A). In some embodiments, the visual representation of the CGR content 7008 obscures a portion of the CGR content (e.g., through blurring, distortion, etc.) and simply conveys the sense of change and color or gradation of the CGR content. As shown in the right portion of Figure 7A, the representation 7006 of a portion of the first user 7202 optionally includes a camera view of the first user's face, or a graphic representation generated based on a camera view of the first user's face. In some embodiments, the representation of a portion of the first user 7202 optionally includes a camera view of the first user's eyes, or a graphic representation generated based on the camera view of the first user's eyes in front of the display surface of the first display generation component.In some embodiments, a representation of a portion of the first user 7202 is displayed on a different display layer from the display layer(s) of one or more graphic elements representing the state of the CGR content displayed via the first display generation component. In some embodiments, the simultaneous display of the representation of the state of the CGR content and the representation of a portion of the first user by the second display generation component provides an indication that the CGR content is displayed via the first display generation component in mixed reality mode and that the first user is provided with a view of the physical environment along with the CGR content. While the first user's face and / or eyes are obscured by the presence of the first display generation component (and optionally the presence of the second display generation component) (e.g., by the presence of the HMD including internal and external displays), the second display generation component displays a visual representation of the first user's face and / or eyes, as well as a representation of the state of the CGR content being viewed by the first user, thereby providing other users in the physical environment surrounding the first user with more information to initiate or refrain from interacting with the first user, or to act appropriately in the presence of the first user.

[0120] Figure 7B follows Figure 7A, showing that the CGR content subsequently progresses further on the first display generation component (e.g., display 7100 or the internal display of the HMD), resulting in a change in the appearance of the first user. For example, the change in the appearance of the first user is due to movement of at least a portion of the first user 7202 (e.g., the first user's eyes or face) relative to the first display generation component within position A7000-a (e.g., while the first user is still wearing the HMD and / or facing the internal display of the HMD) (e.g., movement includes lateral movement of the first user's eyeballs, blinking of the user's eyes, opening and closing of the user's eyes, vertical movement of the user's eyeballs, and / or movement includes movement of the user's face or head relative to the display surface of the first display generation component (e.g., movement away from or toward the first display generation component)). At this point, the CGR content 7002 is still displayed in mixed reality mode, and the representation 7010 (shown, for example, as 7010-b) of the physical environment (e.g., position B including the second user 7204) remains displayed simultaneously with the CGR content 7002 via the first display generation component. In some embodiments, any change in the appearance of the physical environment (e.g., movement of the second user 7204 relative to the first display generation component, the second display generation component, and / or the first user) is similarly reflected in the representation 7010 of the physical environment shown by the first display generation component. In some embodiments, in accordance with movement of a portion of the first user relative to the first display generation component (e.g., movement of the first user's eyes or face), the computing system updates the representation 7006 (shown, for example, as 7006-b in Figure 7B) displayed via the second display generation component 7102.For example, when the first user 7202 or a portion thereof (e.g., the user's face or eyes) moves toward the first edge of the first display generation component 7100 (e.g., the left edge of the display surface of the first display generation component, the upper edge of the display surface of the first display generation component, etc.), the representation 7006 of the portion of the first user shown on the display surface of the second display generation component also moves toward the corresponding second edge of the second display generation component (e.g., the right edge of the display surface of the second display generation component (e.g., corresponding to the left edge of the display surface of the first display generation component), the upper edge of the display surface of the second display generation component (e.g., corresponding to the upper edge of the display surface of the first display generation component), etc.). In addition to updating the representation 7006 of the portion of the first user, the computing system also updates the representation of the state of the CGR content on the second display generation component. For example, the progress bar 7004 (e.g., shown as 7004-b in Figure 7B) is updated to indicate that the playback of the CGR content has progressed by a first amount from the time shown in Figure 7A. In some embodiments, a representation 7008 of the CGR content (shown, for example, as 7008-b in Figure 7B), such as one shown on a second display generation component (e.g., display 7102, an external display of the HMD), is also updated according to the current appearance of the CGR content 7002 shown on a first display generation component (e.g., display 7100, an internal display of the HMD). In some embodiments, by showing real-time updates of the appearance of a portion of the first user (e.g., showing changes and movements of the first user's face and eyes behind the first display generation component) and real-time or periodic updates of the state of the CGR content shown by the first display generation component, it becomes possible to obtain information about the first user's attention state and whether it is appropriate for others in the physical environment surrounding the first user (e.g., at location B) to engage with or interrupt the first user at that moment.In some embodiments, while changes in the appearance of the first user and CGR content are reflected by updates to state information indicated by the second display generation component, any changes in the appearance of the physical environment (e.g., the movement of the second user 7204 relative to the first display generation component, the second display generation component, and / or the first user) are similarly reflected by the representation 7010 of the physical environment indicated by the first display generation component.

[0121] Figure 7C follows Figure 7A, showing that the CGR content then progresses further on the first display generation component 7100, and the second user 7204 moves toward the second display generation component 7204, viewing the second display generation component from a different angle compared to the scenario shown in Figure 7A. At this point, the CGR content 7002 is still displayed in mixed reality mode, and the representation 7010 (shown as, for example, 7010-c) of the physical environment (e.g., position B, including the second user 7204) remains simultaneously displayed via the first display generation component 7100. In accordance with the movement of the second user 7204 toward the second display generation component (and, if the first and second display generation components have a fixed spatial relationship with respect to each other (e.g., fixed back-to-back within the same housing of the HMD), toward the first display generation component, the computing system updates the representation 7010 displayed via the first display generation component (e.g., the display 7100 or the internal display of the HMD) (e.g., shown as 7010-c in Figure 7C). For example, when the second user 7204 or a portion thereof moves toward the third edge (e.g., the right edge, the top edge, etc.) of the display surface of the second display generation component 7102, the representation 7010 of the portion of the second user shown on the display surface of the first display generation component 7100 also moves toward the corresponding fourth edge of the first display generation component (e.g., the left edge of the display surface of the first display generation component (e.g., corresponding to the right edge of the display surface of the second display generation component), the top edge of the display surface of the first display generation component (e.g., corresponding to the top edge of the display surface of the second display generation component)). In accordance with the change in CGR content shown on the first display generation component (e.g., display 7100, internal display of the HMD, etc.), the computing system also updates the representation of the state of the CGR content shown on the second display generation component (e.g., display 7102, external display of the HMD, etc.).For example, the progress bar 7004 (shown, for example, as 7004-c in Figure 7C) is updated to indicate that playback of the CGR content has progressed by a second amount from the time shown in Figure 7A and a third amount from the time shown in Figure 7B. In some embodiments, a representation 7008 of the CGR content 7002 (shown, for example, as 7008-c in Figure 7C), such as one shown on a second display generation component (e.g., display 7102, an external display of the HMD), is also updated according to the current appearance of the CGR content 7002 shown on a first display generation component (e.g., display 7100, an internal display of the HMD). The contrast in appearance of the state information shown in Figures 7A and 7B (including, for example, representation 7006, representation 7008, and progress bar 7004) indicates that, for the same relative spatial position between the first display generation component and the portion of the first user 7202 represented by the state information shown by the second display generation component, representation 7006 of the portion of the first user 7202 is displayed at a different depth than that of representation 7008 of the CGR content, and optionally at a different depth than that of other state information (e.g., progress bar 7004). The difference in display depth from the display surface of the second display generation component 7102 or from the position of the second user 7204 results in a visual parallax effect. For example, when a second user 7204 moves to a second display generation component (e.g., a display 7102, an external display of an HMD, etc.), the representation 7006 of a portion of the first user 7202 and the representation 7008 of the CGR content appear to move by different amounts on the display surface of the second display generation component (and appear to move relative to each other). In some embodiments, the representation 7008 of the CGR content is displayed as a diffusion layer between the representation 7006 of a portion of the first user and other representations of state information (e.g., the title of the CGR content, a progress bar 7004, etc.). In some embodiments, the representation of a portion of the first user is displayed in the display layer furthest from the display surface of the second display generation component, compared to the display layers of the representation of the CGR content and other representations of state information shown by the second display generation component.

[0122] In some embodiments, the first and second display generation components are positioned back-to-back within an HMD worn on the head of a first user or positioned in front of the user's face (for example, with their respective display surfaces facing in different directions (e.g., substantially opposite directions)). In some embodiments, the second display generation component displays a visual representation of the first user's eyes, generated based on an actual image of the first user's eyes using one or more image processing filters. For example, the visual representation of the first user's eyes is optionally generated by reducing the opacity of the camera image of the first user's eyes, increasing the transparency, reducing the color saturation level, reducing the luminance level, reducing the pixel resolution, reducing the color resolution, etc. In some embodiments, the amount of modification applied to various display characteristics of the individual camera images of the first user's eyes is optionally specified for various display characteristic values ​​of the representation 7008 of the CGR content simultaneously displayed by the second display generation component 7102. For example, if the representation of CGR content is relatively dark (e.g., having a first range of luminance values), the representation of the eyes is also made darker, more translucent, and / or less color-saturated (e.g., having a second range of luminance values, a second range of transparency values, and a second range of color-saturation values ​​selected based on the first range of luminance values); if the representation of CGR content is brighter (e.g., having a second range of luminance values ​​greater than the first range of luminance values), the representation of the eyes is made brighter, less translucent, and / or more color-saturated (e.g., a third range of luminance values, a third range of transparency values, and a third range of color-saturation values ​​selected based on the second range of luminance values). In some embodiments, other display characteristics (e.g., color-saturation, pixel resolution, color resolution, gradation, etc.) are used as a basis for selecting value ranges for the display characteristics of representations of parts of the user (e.g., the user's face or eyes).In some embodiments, a representation of the first user's eye is generated by applying one or more pre-configured image filters, such as a blur filter, a color filter, or a luminance filter, which alter the original appearance of the first user's eye when the representation is displayed by a second display generation component.

[0123] In some embodiments, a representation of the CGR content (e.g., representation 7008) shown by a second display generation component is generated by applying a diffuse filter to the CGR content displayed by the first display generation component (e.g., all visible content, media content only, or optionally, visible content excluding pass-through views of the physical environment). For example, the color and gradation of the scene are preserved in the representation 7008 of the CGR content, but the contours of objects within the CGR content are blurred and not clearly defined within the representation 7008 of the CGR content. In some embodiments, the representation of the CGR content is semi-transparent, through which a representation 7006 of a portion of the first user is visible. In some embodiments, graphical user interface elements representing metadata associated with the CGR content (e.g., progress bar 7004, title of the CGR content, etc.) are displayed by the second display generation component (e.g., on the same or a different display layer as the representation 7008 of the CGR content, and / or on the same or a different display layer as the representation 7006 of a portion of the first user). In some embodiments, graphical user interface elements representing metadata associated with CGR content are displayed with higher pixel resolution, higher color resolution, higher color saturation, higher opacity, higher luminance, and / or better defined contours compared to the representation of the CGR content 7008 and / or the representation of a portion of the first user 7006.

[0124] In some embodiments, a portion of the first user (e.g., the first user's face or eyes) moves relative to a first display generation component (e.g., display 7100, an internal display of the HMD), while the CGR content 7002 presented by the first display generation component remains unchanged. In such cases, the representation 7006 of the user's portion is optionally updated on the second display generation component 7102 without updating the representation 7006 of the CGR content and the progress bar 7004. In some embodiments, the CGR content is not displayed or is paused, and the first user is viewing a pass-through view of the physical environment through the first display generation component without simultaneous display of the CGR content, and the second display generation component optionally updates the representation of the first user's portion in accordance with changes in the appearance of the first user's portion (e.g., due to movement or other changes of the user's portion), without displaying any representation of the CGR content or showing a static or paused representation of the CGR content.

[0125] In some embodiments, the CGR content changes on the first display generation component, while a portion of the first user does not change its appearance (e.g., it does not move or change for other reasons). Thus, the representation 7006 of the portion of the first user remains unchanged, and the second display generation component updates only the representation 7008 of the CGR content and other indicators of the state of the CGR content (e.g., progress bar 7004) in accordance with the changes in the CGR content shown by the first display generation component.

[0126] In some embodiments, if changes in both the CGR content and the appearance of a portion of the first user are detected during the same period (e.g., simultaneously and / or between pre-set time windows of each other), the second display generation component updates both the visual representation of the portion of the user and one or more graphic elements (e.g., the representation of the CGR content 7008 and the progress bar 7004) indicating the state of the CGR content, in accordance with the detected changes.

[0127] Figures 7A to 7C illustrate that, via the second display generation component 7102, the appearance of the visual representation (e.g., representation 7006) of the first user 7202 (e.g., the first user's eyes or face) is updated in accordance with changes in the first user's appearance (e.g., changes in the first user's eye movements or facial expressions, changes in lighting, etc.), and via the second display generation component 7102, one or more graphic elements (e.g., progress bar 7004, representation 7008 of CGR content 7002, etc.) that provide visual indication of content within the CGR environment shown via the first display generation component 7100 are updated in accordance with changes in the CGR environment. In the exemplary scenarios shown in Figures 7A to 7C, the immersion level associated with the CGR content and the attention state of the first user 7202 remain unchanged and correspond to the intermediate immersion level associated with the presentation of the CGR content. In some embodiments, the level of immersion associated with the presentation of CGR content and the corresponding attention state of the first user are optionally changed over a period of time, for example, increasing to a higher level of immersion and a more engaged user attention state, or decreasing to a lower level of immersion and a less engaged user attention state. In some embodiments, state information presented by a second display generation component is updated based on the change in the level of immersion to which the CGR content is presented by the first display generation component. In some embodiments, updates to the state information include updates to the representation of a portion of the first user (e.g., updates to the visibility of the representation of a portion of the first user, updates to the appearance of the representation of a portion of the first user, etc.).

[0128] In some embodiments, the computing system is configured to display CGR content 7002 at at least a first, second, and third level of immersion. In some embodiments, the computing system transitions the CGR content displayed via the first display generation component between different levels of immersion in response to a series of events (e.g., the natural termination or progression of an application or experience, the start, stop, and / or pause of an experience in response to user input, a change in the level of immersion of an experience in response to user input, a change in the state of a computing device, a change in the external environment, etc.). In some embodiments, the first, second, and third levels of immersion correspond to increasing the amount of virtual content present in the CGR environment and / or decreasing the amount of representation of the surrounding physical environment present in the CGR environment (e.g., a representation of a portion of the physical environment in front of position B or the display surface of the second display generation component 7102). In some embodiments, the first, second, and third immersion levels correspond to different modes of content display having increased image fidelity (e.g., increased pixel resolution, increased color resolution, increased color saturation, increased luminance, increased opacity, increased image detail, etc.) and / or spatial range (e.g., angular range, spatial depth, etc.) of computer-generated content, and / or decreased image fidelity and / or spatial range of the representation of the surrounding physical environment (e.g., representation of a portion of the physical environment in front of the display surface of position B or the second display-generating component). In some embodiments, the first immersion level is a pass-through mode in which the physical environment (e.g., a portion of the physical environment in front of the display surface of position B or the second display-generating component) is fully visible to the first user through the first display-generating component (e.g., as a camera view of the physical environment or through the transparent portion of the first display-generating component).In some embodiments, the CGR content presented in pass-through mode includes a pass-through view of the physical environment, which has the minimum amount of virtual elements simultaneously visible as a view of the physical environment, or only virtual elements surrounding the user's view of the physical environment (e.g., indicators and controls displayed in the peripheral area of ​​the display). Figure 7E shows examples of a first level of immersion associated with CGR content 7002 according to some embodiments. For example, the view of the physical environment (e.g., a portion of the physical environment in front of the display surface of the second display generation component (e.g., also a portion of the physical environment in front of the first user)) occupies the central and majority areas of the field of view provided by the first display generation component, and only a few controls (e.g., movie title, progress bar, playback controls (e.g., play button)) are displayed in the peripheral area of ​​the field of view provided by the first display generation component. In some embodiments, the second level of immersion is a mixed reality mode in which a pass-through view of the physical environment is augmented with virtual elements generated by a computing system, and the virtual elements occupy the central and / or majority of the user's field of view (e.g., the virtual content is integrated with the physical environment within the view of the computer-generated environment). Examples of the second level of immersion associated with CGR content 7002 according to some embodiments are shown in Figures 7A-7C. In some embodiments, the third level of immersion is a virtual reality mode in which the user's view of the physical environment is completely replaced or blocked by a view of virtual content provided by the first display generation component. Figure 7D shows an example of the third level of immersion associated with CGR content 7002 according to some embodiments.

[0129] As shown in Figure 7D following Figure 7C, the computing system has a switch from displaying CGR content 7002 in mixed reality mode to displaying CGR content 7002 (e.g., movie X, shown as 7002-d) in virtual reality mode without representation of the physical environment (e.g., position B including the second user 7204 (e.g., a portion of the physical environment in front of the display surface of the second display generation component 7102)). In some embodiments, the switch performed by the computing system is in response to a request from the first user (e.g., a gesture input that meets a pre-set criterion for changing the level of immersion of the CGR content (e.g., the first user raising their hand away from the HMD)). In conjunction with the switch from displaying CGR content 7002 in mixed reality mode via the first display generation component 7100 to displaying CGR content 7002 in virtual reality mode, the computing system changes the state information displayed via the second display generation component 7102. As shown in Figure 7D, one or more graphic elements representing the CGR content 7002 (e.g., a title, a progress bar 7004 (e.g., shown as 7004-d in Figure 7D), and a representation 7008 (e.g., shown as 7008-d in Figure 7D)) are still displayed and continue to update in accordance with changes in the CGR content 7002 (e.g., shown as 7002-d in Figure 7D) represented by the first display generation component 7100, but the representation 7006 of a part of the first user (e.g., the eyes or face of the first user) is no longer displayed by the second display generation component 7100. In some embodiments, instead of completely ceasing to display a portion of the first user's representation, the computing system reduces the visibility of the portion of the first user's representation to the visibility of other state information on the second display generation component (e.g., representation of CGR content, representation of metadata related to CGR content or the user) (e.g., by reducing luminance, reducing color resolution, reducing opacity, reducing pixel resolution, etc.).In some embodiments, a representation of a portion of the first user 7006 is optionally displayed with reduced visibility relative to its previous appearance (e.g., completely invisible, or with reduced luminance, increased transparency, reduced opacity, reduced color saturation, increased blur level, etc.) to indicate an increased level of immersion associated with the CGR content represented by the first display generation component. Other state information (e.g., the representation of the CGR content 7008 and the progress bar 7004) is updated continuously or periodically in accordance with changes in the CGR content 7002 represented by the first display generation component, while such other state information remains displayed by the second display generation component 7102 (e.g., without a reduction in visibility relative to its previous appearance, unless the reduction is due to a change in the appearance of the CGR content).

[0130] In some embodiments, the switch from mixed reality mode to virtual reality mode is triggered by the movement of the second user 7204 out of the estimated field of view that the first user would have had if the first user's eyes were not obstructed by the presence of the first and / or second display generation components. In some embodiments, the switch from mixed reality mode to virtual reality mode is triggered by the movement of the second user 7204 out of the physical environment surrounding the first user (e.g., out of the room occupied by the first user). In some embodiments, the computing system stops displaying a representation of the physical environment (e.g., a representation of location B (e.g., a representation of a portion of the physical environment in front of the first user)) if there are no other users present in the physical environment. In some embodiments, the estimated field of view that the first user would have if the view of the first user at position B were not obstructed by the presence of the first and / or second display generation components, as well as the movement of the second user 7204 into the physical environment surrounding the first user (e.g., a room occupied by the first user), a default gesture performed by the second user (e.g., the second user waving to the first user), the second user moving within a threshold distance range of the first user, etc., are optionally used as conditions to trigger a switch from virtual reality mode to mixed reality mode. In some embodiments, in conjunction with switching the display mode from virtual reality mode to mixed reality mode, the computing system restores the level of visibility of the representation 7006 of a portion of the first user among the elements of state information shown by the second display generation component 7102 (for example, restoring the visibility of the representation of a portion of the first user if the representation was not visible, or increasing the luminance, color saturation, pixel resolution, opacity, and / or color resolution of the representation of a portion of the user, etc.).Accordingly, in mixed reality mode, the first display generation component (e.g., display 7100, an internal display of the HMD, etc.) displays a representation (e.g., representation 7010) of a portion of the physical environment in front of the display surface of the second display generation component (and accordingly, in front of the first user if the first and second display generation components are surrounded back-to-back within the same housing of the HMD worn by the first user), along with computer-generated virtual content (e.g., movie X).

[0131] As shown in Figure 7E (for example, after Figure 7C or Figure 7D), the computing system switches from displaying CGR content in mixed reality mode (for example, as shown in Figures 7A to 7C) or virtual reality mode (for example, as shown in Figure 7D) to displaying CGR content (for example, shown as 7010-e in Figure 7E) in a representation of the physical environment (for example, position B including a second user 7204 (for example, a portion of the physical environment in front of the display surface of the second display generation component 7102)) on the peripheral area of ​​the display (for example, the upper and lower edges of the display), and optionally in a full pass-through mode or reality mode (for example, when movie X is finished and no longer shown) with only a minimum amount of virtual content (for example, only indicators (for example, the title of movie X, progress bar, playback controls, etc.) and no playback content). In some embodiments, the display mode switching is performed by the computing system in response to the end or pause of playback of the CGR content and / or a request from the first user (e.g., a gesture input that meets a pre-set criterion for changing the immersion level of the CGR content, such as placing a hand above the first user's eyebrows or pinching down on the surface of the HMD). In some embodiments, in conjunction with the switch from displaying the CGR content in mixed reality mode or virtual reality mode via the first display generation component 7100 to displaying the CGR content in full passthrough mode or reality mode, the computing system changes the state information displayed via the second display generation component 7102.As shown in Figure 7E, one or more graphic elements indicating the appearance and state of the CGR content (e.g., a title, a progress bar 7004 (e.g., shown as 7004-c in Figure 7C and 7004-d in Figure 7D, etc.), and a representation 7008 (e.g., shown as 7008-c in Figure 7C and 7008-d in Figure 7D, etc.)) are no longer displayed by the second display generation component, and a representation 7006 of a portion of the first user (e.g., shown as 7006-e in Figure 7E) is fully displayed with increased visibility by the second display generation component (e.g., it becomes visible where it was previously invisible, or it is displayed with increased luminance, decreased transparency, increased opacity, increased color saturation, increased pixel resolution, increased color resolution, reduced blur level, etc.). In some embodiments, the representation 7006 of a portion of the first user is optionally displayed with increased visibility compared to its previous appearance to indicate a decrease in the level of immersion associated with the CGR content shown by the first display generation component. In some embodiments, the representation 7006 of a portion of the first user is continuously updated in accordance with changes in the appearance of the first user 7202 while the representation 7006 of a portion of the first user is displayed by the second display generation component.

[0132] In some embodiments, the switch from mixed reality or virtual reality mode to full passthrough mode or reality mode is triggered by the movement of the second user 7204 into an estimated field of view that the first user would have had if the first user's eyes were not obstructed by the presence of the first and / or second display generation components. In some embodiments, the switch from mixed reality or virtual reality mode to full passthrough mode or reality mode is triggered by the movement of the second user 7204 into personal space within a threshold distance from the first user 7202 (e.g., within arm's length from the first user, within 3 feet from the first user, etc.). In some embodiments, when the computing system enters full passthrough mode or reality mode (e.g., when a pre-set condition is met, e.g., when a pre-set person (e.g., spouse, teacher, teammate, child, etc.) enters the estimated field of view of the first user 7202), it stops displaying CGR content via the first display generation component and displays only a representation of the physical environment (e.g., location B, the physical environment in front of the first user, etc.). In some embodiments, the movement of a second user 7204 out of the estimated field of view of a first user 7202 and / or out of personal space within a threshold distance from the first user 7202, and / or other conditions, are used to trigger an automatic switch back from full pass-through mode or reality mode to mixed reality mode or virtual reality mode (e.g., a pre-set mode or a previous mode). In some embodiments, in conjunction with switching the display mode from full pass-through mode to virtual reality mode or mixed reality mode, the computing system restores the level of visibility of the representation 7006 of a portion of the state information elements shown by the second display generation component 7102 (e.g., by reducing its visibility without completely stopping displaying it or making it completely invisible), and restores the level of visibility of the representation 7008 of the CGR content (e.g., by increasing its visibility).

[0133] Figures 7C to 7E illustrate, in some embodiments, transitions from a second level of immersion (e.g., mixed reality mode) down to a first level of immersion (e.g., pass-through mode) and up to a third level of immersion (e.g., virtual reality mode), as well as the corresponding changes in information displayed by the first display generation component 7100 and the second display generation component 7102. In some embodiments, direct transitions between any two of the three levels of immersion are possible depending on different events that satisfy the respective criteria for such direct transitions. Accordingly, the information displayed by the first display generation component and the second display generation component are updated (e.g., changing the visibility of different components of the information (e.g., CGR content, representation of the physical environment, representation of the CGR content, representation of a part of the first user, representation of static metadata associated with the CGR content, etc.)) to reflect the current level of immersion in which the CGR content is displayed by the first display generation component 7100.

[0134] In some embodiments, as shown in Figure 7E, a representation of a portion of the first user 7006 (e.g., a representation of the first user's face or eyes) is displayed without simultaneous display of the representation of the CGR content (e.g., without overlaying diffused versions of the CGR content, title, or progress bar) or with increased visibility relative to the representation of the CGR content (e.g., the visibility of representation 7006 is increased relative to its previous level, the visibility of the representation of the CGR content is decreased relative to its previous level, and / or some of the graphic elements for representing the CGR content are no longer displayed, etc.).

[0135] In some embodiments, a representation 7006 of a portion of the first user (e.g., a representation of the first user's face or eyes) is displayed together with a representation 7008 of the CGR content (e.g., together with a superimposed diffused version of the CGR content) (e.g., with equivalent visibility to the representation 7008 of the CGR content (e.g., increased or decreased visibility of representation 7006 and / or representation 7008 relative to their respective previous levels)) as a result of the computing system switching from displaying the CGR content using virtual reality mode or pass-through mode to displaying the CGR content using mixed reality mode.

[0136] In some embodiments, when a computing system switches from displaying CGR content using mixed reality mode to displaying CGR content using virtual reality mode, a representation of a part of the first user 7006 (e.g., a representation of the first user's face or eyes) is not displayed together with the representation of the CGR content 7008 (e.g., not displayed together with the diffused version of the CGR content), or is displayed with reduced visibility relative to the representation of the CGR content 7008.

[0137] In some embodiments, the computing system may display CGR content using other special display modes such as private mode, silent (Do-Not-Disturb, DND) mode (DND mode), and parental control mode. When one or more of these special display modes are turned on, the way in which state information is displayed and / or updated on the second display generation component is adjusted from the way in which state information is displayed and / or updated on the second display generation component when such special modes are not turned on (for example, the method described above with reference to Figures 7A to 7E).

[0138] For example, Private Mode is optionally activated by a computing system or a first user to hide state information associated with the CGR content currently displayed by the first display generation component, and / or state information associated with the attention state of the first user. In some embodiments, while Private Mode is turned on (e.g., at the request of the first user), the representation 7006 of a portion of the first user and / or the representation 7008 of the CGR content are no longer updated, cease to be displayed, and / or are replaced with other placeholder content on the second display generation component, and as a result no longer reflect changes detected in the appearance of a portion of the first user, and / or changes detected in the CGR content displayed by the first display generation component. In some embodiments, Private Mode is activated in response to a user request (e.g., a pre-configured gesture input, a pre-configured voice command, etc.) detected by the computing system (e.g., when the computing system is displaying the CGR content to the first user using mixed reality mode or virtual reality mode, and / or before the CGR content is started, etc.). In some embodiments, private mode is activated in response to a user accessing specific CGR content associated with a pre-configured privacy level that exceeds a first threshold privacy level (e.g., a default privacy level, a privacy level associated with a first user). In some embodiments, while private mode is enabled, the representation 7006 of the first user and / or the representation 7008 of the CGR content are no longer updated, cease to be displayed, and / or are replaced with other placeholder content, and as a result no longer reflect changes in the level of immersion to which the CGR content is displayed by the first display generation component.Private mode allows the first user to enjoy greater privacy and to share less information about their attention state, level of immersion, and the content they are viewing using the first display generation component through the content displayed by the second display generation component.

[0139] In some embodiments, DND mode is automatically turned on by the computing system based on conditions set in advance by the first user and / or pre-configured by the computing system to indicate to the external environment that the first user does not want their engagement with CGR content to be interrupted or interfered with by others in the external environment. In some embodiments, DND mode is optionally applicable to other intrusion events occurring within the computing system and / or the surrounding environment. For example, in some embodiments, in response to the activation of DND mode, the computing system optionally activates noise cancellation to block out sounds from the surrounding environment, stops / pauses the presentation of notifications and / or warnings on the first display generation component, reduces the intrusiveness of how notifications and / or warnings are presented in the CGR environment shown by the first display generation component (e.g., select visual warnings instead of audio warnings, select short warning sounds instead of voice output, reduce the visual prominence of notifications and warnings, etc.), automatically transfers calls to voicemail and / or displays a silent mode sign on the second display generation component without notifying the first user, etc. In some embodiments, one or more methods used by the computing system to reduce the inconvenience of events to the first user involve changing how representations of the physical environment (e.g., representation 7010, representation of location B, representation of a portion of the physical environment in front of the first user, etc.) are displayed on a first display generation component, and / or changing how state information is displayed by a second display generation component. In some embodiments, DND mode is optionally turned on while the computing system is displaying CGR content using mixed reality mode or virtual reality mode. In some embodiments, in response to DND mode being turned on, the computing system optionally displays a visual indicator (e.g., a text label "DND" on the external display of the HMD, a red border lit around the external display of the HMD, etc.) via a second display generation component to indicate that DND mode is active.In some embodiments, while DND mode is active on the computing system, the representation of CGR content is optionally updated in accordance with changes in the CGR content displayed by the first display generation component, but the representation of a portion of the first user is no longer updated, replaced by placeholder content, or ceases to be displayed by the second display generation component (e.g., regardless of changes in the appearance of a portion of the first user (e.g., changes in the first user's eyes) and / or changes in the level of immersion to which the CGR content is displayed by the first display generation component).

[0140] In some embodiments, parental mode is turned on to override the normal display of state information by the second display generation component (as described, for example, with reference to Figures 7A–7E). Parental mode is turned on so that a parent, teacher, or supervisor can view and monitor the CGR content presented to the first user, and optionally, inputs provided by the first user to modify and / or interact with the CGR content. In some embodiments, parental mode is optionally turned on by the second user while the CGR content is being presented by the first display generation component (e.g., via pre-configured gesture input, touch input on the second display generation component or the HMD housing, voice commands, etc.). In some embodiments, parental mode is optionally turned on before specific CGR content is initiated on the first display generation component (e.g., via interaction with a user interface presented by the first display generation component, interaction with a user interface presented by the second display generation component, interaction with the computing system housing or other input devices, etc.) and remains turned on while the specific CGR content is displayed by the first display generation component. In some embodiments, while parental mode is turned on, the computing system displays the same CGR content simultaneously through both the first and second display generation components, regardless of changes in immersion level and / or whether private mode is turned on. In some embodiments, the computing system displays only the virtual content portion of the CGR content presented by the first display generation component on the second display generation component.In some embodiments, while Parental Mode is enabled, the computing system does not display the representation 7006 of a portion of the first user as part of state information shown using the second display generation component (e.g., if Parental Mode is used simply to monitor content shown to the first user and not to monitor the first user themselves). In some embodiments, while Parental Mode is enabled, the computing system displays the representation 7006 of a portion of the user and the CGR content with equivalent visibility to the CGR content (e.g., the visibility of the representation 7006 is improved compared to the previous level of visibility it had when Parental Mode was not enabled) (e.g., if Parental Mode is used to monitor content shown to the first user, as well as the first user's attention state and appearance). In some embodiments, whether the representation of a portion of the user is displayed by the second display generation component during Parental Mode is determined according to how Parental Mode is activated (e.g., using a first type of input versus using a second type of input, using a first control versus using a second control, etc.). In some embodiments, whether a portion of the user's representation is displayed by the second display generation component during parental mode is determined by whether private mode is enabled. For example, if private mode is enabled, the portion of the user's representation is not displayed by the second display generation component along with the CGR content; if private mode is disabled, the portion of the user's representation is displayed by the second display generation component along with the CGR content.In some embodiments, while parental mode is enabled, changes in the immersion level at which the first display generation component displays CGR content do not alter the information displayed by the second display generation component (for example, the same CGR content may optionally still be displayed to both the first and second display generation components, either without any change in the current visibility level of the representation of a portion of the first user, or without any representation of the first user).

[0141] In some embodiments, the visibility and information density of state information displayed by the second display generation component are dynamically adjusted by the computing system according to the distance of the second user who is positioned (e.g., directly or partially in front of the display surface of the second display generation component) in a location that enables the second user to view the content displayed by the second display generation component. For example, when the second user moves closer to the display surface of the second display generation component (e.g., moves within a threshold distance, moves within a threshold field of view, etc.) (e.g., moves closer to the first user and the first display generation component if the first and second display generation components are positioned back-to-back within the same housing of an HMD worn by the first user), the computing system changes (e.g., increases) the amount of information detail provided on the second display generation component (e.g., graphic feature detail, amount of text characters per unit display area, color resolution, pixel resolution, etc.) to inform the second user of the state of the first user and the state and metadata of the CGR content. Accordingly, when a second user moves far away from the display surface of the second display generation component (e.g., moves beyond a threshold distance, moves outside the threshold viewing angle, etc.), the computing system changes the amount of information detail provided on the second display generation component in the opposite direction (e.g., reduces the amount of information detail).

[0142] In some embodiments, the computing system automatically transitions from displaying the computer-generated experience in a fully immersive mode (e.g., displaying a virtual reality environment or displaying CGR content at a third level of immersion) to displaying the computer-generated experience in a less immersive mode (e.g., displaying indications of the physical environment within the virtual reality environment (e.g., displaying the outlines of people and objects in the physical environment as visual distortions, shadows, etc.), displaying pass-through portions (e.g., camera views of the physical environment) in the view of the computer-generated environment) in response to detecting changes in the surrounding physical environment that meet pre-set criteria (e.g., people entering a room or coming within a threshold distance of a first user, another user waving or making a gesture towards the first user, etc.). In some embodiments, in conjunction with automatically changing the level of immersion of the computer-generated environment displayed via the first display generation component, the computing system also modifies state information displayed via the second display generation component, including increasing the visibility of the visual representation of a portion of the user of the computing system (for example, increasing the visibility of the user's visual representation includes switching from not displaying the visual representation of a portion of the user to displaying it, or increasing the luminance, sharpness, opacity, and / or resolution of the visual representation of a portion of the user). In this way, the visual barrier separating the first user from others in the surrounding environment (for example, the presence of a display generation component on the first user's face) is simultaneously reduced, facilitating more informative interactions between the first user and surrounding users.In some embodiments, if the computing system reduces the level of immersion of the content displayed on the first display generation component in response to an action by the second user (for example, in response to the second user waving to the first user, and / or in response to the second user moving too close to the first user, the computing system may stop displaying the representation of the CGR content, or not display the representation of the CGR content at all, and instead display only a representation of a portion of the first user (for example, the first user's face or eyes) on the second display generation component (for example, to inform the second user that the first user can see the second user through the first display generation component). In some embodiments, if the computing system increases the level of immersion of the content displayed on the first display generation component in response to an action by a second user (e.g., in response to the second user putting on an HMD and / or in response to the second user moving away from the first user, the computing system redisplays the representation of the CGR content on the second display generation component and stops displaying (or reduces the luminance, sharpness, opacity, color, and pixel resolution, etc., of a representation of a part of the first user (e.g., the first user's face or eyes)).

[0143] Further details regarding the user interface and operating modes of the computing system are provided by referring to Figures 7F to 7Q, Figures 8 to 13, and the accompanying descriptions.

[0144] Figures 7F to 7I show a computing system (e.g., computing system 101 in Figure 1 or computing system 140 in Figure 4) that includes at least a first display generation component (e.g., display 7100, an internal display of the HMD) and a second display generation component (e.g., display 7102, an external display of the HMD), wherein the first display generation component of the computing system is configured to display visual content (e.g., a user interface, computer-generated experience, media content, etc.) to a user (e.g., first user 7202) when the computing system determines that the first display generation component is positioned in a predetermined spatial relationship (e.g., in a predetermined orientation with respect to a user (e.g., first user 7202) or a part of the user (e.g., the face or eyes of the first user) (e.g., the display surface facing the face or eyes of the first user). In some embodiments, before displaying visual content to the user via a first display generation component (e.g., display 7100, an internal display of the HMD), the computing system uses a second display generation component (e.g., display 7102, an external display of the HMD) to display one or more user interface elements (e.g., a user interface, a computer-generated experience (e.g., AR content, VR content, etc.), media content, etc.) that prompt the user about visual content available to be displayed via the first display generation component when conditions for triggering the display of content (e.g., conditions relating to the spatial relationship between the first display generation component and the user) are met. In some embodiments, one or more user interface elements include user interface objects that convey contextual information (e.g., current time, location, external conditions, user identification information, and / or the current state of the computing system, or other prompts, notifications, etc.) based on which available content is now available to be displayed via the first display generation component.Figures 7F-7G and 7H-7I show two parallel embodiments illustrating that different computer-generated experiences are displayed by the first display generation component based on different states of the computing system, which are reflected by different user interface objects shown by the second display generation component, before a pre-defined spatial relationship between the first display generation component and the user is satisfied.

[0145] As shown in the left-hand portions of Figures 7F and 7H, the first display generation component (e.g., display 7100) is located at position A7000-a with no user facing the display surface of the first display generation component. As a result, the first display generation component is not currently displaying any CGR content. As shown in the right-hand portions of Figures 7F and 7H, the second display generation component (e.g., display 7102) is located at position B7000-b and displays one or more user interface elements, each including a first user interface element (e.g., circle 7012) corresponding to a first computer-generated experience 7024 available for display via the first display generation component, given the current context state of the computing system shown in Figure 7F, and a second user interface element (e.g., square 7026) corresponding to a second computer-generated experience 7030 available for display via the first display generation component, given the current context state of the computing system shown in Figure 7H. The contextual state of the computing system is determined based on, for example, the current time, the current position of the first display generation component (which in some embodiments is also the position of the first user and the second display generation component), the user identification information or authorization level of the first user 7202 present in front of the display surface of the second display generation component, the reception or generation of notifications for individual applications by the computing system, the occurrence of pre-configured events on the computing system, and / or other contextual information.

[0146] In the exemplary scenarios shown in Figures 7F and 7H, the first user 7202 is shown to be located at position B7000-b. In some embodiments, the spatial relationship between the second display generation component (e.g., display 7102, an external display of the HMD) and the first user 7202 is such that the first user 7202 is positioned to view one or more user interface elements (e.g., user interface element 7012 and user interface element 7026, respectively) presented by the second display generation component. For example, the first user 7202 is facing the display surface of the second display generation component when one or more user interface elements are displayed. In some embodiments, the second display generation component is an external display of the HMD, which also includes an internal display (e.g., the first display generation component represented by display 7100) configured to present CGR content corresponding to user interface elements displayed on the external display of the HMD. In such embodiments, the spatial relationship represented by the simultaneous presence of the display 7102 and the first user 7202 at the same location B7000-b corresponds to the first user being in a portion of the physical environment that the external display of the HMD faces (e.g., the physical environment that also accepts the second display generation component and the first user 7202). In some embodiments, the second display generation component is positioned back-to-back with the first display generation component such that a portion of the physical environment in front of the display surface of the second display generation component 7102 (e.g., location B7000-b including the display surface of the second display generation component) is also within the pass-through view provided by the first display generation component.For example, a physical object 7014 (or both physical objects 7014 and 7028 in Figure 7H) located in part of the physical environment in front of the display surface of the second display generation component 7102 will be within the pass-through view provided by the first display generation component when the first user 7102 moves to the display surface of the first display generation component (e.g., moves to position A7000-a and / or faces the display surface of the first display generation component). In some embodiments, the computing system displays one or more user interface elements only in response to detecting that the first user is in a position to view content displayed by the second display generation component, and stops using the second display generation component to display one or more user interface elements when the user is not in a position relative to the second display generation component that would allow the user to view content presented by the second display generation component. In some embodiments, the computing system displays one or more user interface elements in response to detecting an event that indicates the availability of a first computer-generated experience based on the current state or context of the computing system (e.g., a time-based alert or location-based notification is generated on the computing system, the user takes the HMD out of a bag, the user turns on the HMD, the user plugs the HMD into a charging station, etc.). In some embodiments, the computing system discloses one or more user interface objects on a second display-generating component only when the first display-generating component is not displaying any CGR content (e.g., the first display-generating component is inactive, in a power-saving state, etc.).In some embodiments, the examples in Figures 7F and 7H, which show that the first user 7202 is not simultaneously in a position to view the first display generation component and position A or the content displayed by the first display generation component, correspond to the first user not having the internal display of the HMD in front of their face or eyes (for example, by holding the HMD with the internal display facing the user's face, or by wearing the HMD on the user's head). In some embodiments, one or more user interface objects are displayed on the second display generation component regardless of whether the first user is in a position to view the content shown on the second display generation component.

[0147] As shown in Figure 7G following Figure 7F, and Figure 7I following Figure 7H, while the computing system is displaying one or more user interface objects (e.g., circle 7012 or square 7026, respectively) using the second display generation component, the computing system detects that the first display generation component (e.g., display 7100 or the internal display of the HMD) is currently in a pre-configured spatial relationship with respect to the first user 7202 (e.g., by movement of the first user 7202, movement of the first display generation component, or both). In the embodiments shown in Figures 7F and 7G, the first user 7202 moves to a position A7000-a in front of the display surface of a first display generation component (e.g., display 7100, internal display of the HMD, etc.), and in response to detecting that the first user has moved to a position A7000-a in front of the display surface of the first display generation component, the computing system displays individual computer-generated experiences (e.g., first computer-generated experience 7024 or second computer-generated experience 7030) corresponding to one or more user interface objects (e.g., circle 7012 or square 7026, respectively) previously shown on the second display generation component via the first display generation component. For example, as shown in Figure 7G, the computing system displays a first computer-generated experience 7024 in response to an event indicating that the relative movement of the first display-generating component and the first user has placed the first display-generating component and the first user (or the face or eyes of the first user) into a pre-defined spatial relationship or configuration (for example, the first user is facing the display surface of the first display-generating component, or the HMD is positioned in front of the user's eyes, or the HMD is positioned on the user's head, etc.) while a first user interface object (e.g., circle 7012) is being displayed by a second display-generating component (Figure 7F).As shown in Figure 7I, while a second user interface object (e.g., a square 7026) is displayed by the second display generation component (Figure 7H), the computing system displays a second computer-generated experience 7030 in response to events indicating that the relative movement of the first display generation component and the first user (or the first user's face or eyes) has placed the first display generation component and the first user (or the first user's face or eyes) into a pre-defined spatial relationship or configuration (e.g., the first user is facing the display surface of the first display generation component, or the HMD is positioned in front of the user's eyes, or the HMD is positioned on the user's head, etc.).

[0148] As shown in Figure 7G following Figure 7F, an embodiment of the first computer-generated experience 7024 is an augmented reality experience that represents a portion of the physical environment in front of the display surface of the second display-generating component (e.g., a portion of the physical environment in front of the external display of the HMD, which is also a portion of the physical environment in front of the first user wearing the HMD and / or facing the internal display of the HMD). The first computer-generated experience optionally includes a representation 7014' of the physical object 7014, augmented with some virtual content (e.g., a virtual ball 7020 placed on a representation 7014' of the physical object 7014, and / or some other virtual object). In some embodiments, the first computer-generated experience is a purely virtual experience and does not include a representation of the physical environment surrounding the first display-generating component and / or the second display-generating component.

[0149] As shown in Figure 7I following Figure 7H, an embodiment of the second computer-generated experience 7030 is an augmented reality experience that represents a portion of the physical environment in front of the display surface of the second display-generating component (e.g., a portion of the physical environment in front of the external display of the HMD, which is also a portion of the physical environment in front of the first user wearing the HMD and / or facing the internal display of the HMD). The second computer-generated experience optionally includes physical objects 7014 and 7028, which are stacked together and augmented with some virtual content (e.g., a virtual box 7032 placed next to representations of physical objects 7014 and 7028 7014' and 7028', or some other virtual object). In some embodiments, the second computer-generated experience is a purely virtual experience and does not include representations of the physical environment surrounding the first display-generating component and / or the second display-generating component.

[0150] In some embodiments, as described above in this disclosure, the first display generating component is an internal display of the HMD, and the second display generating component is an external display of the HMD, and the spatial relationship represented by the simultaneous presence of the display 7100 and the first user 7202 at the same location A7000-a corresponds to the first user wearing or holding the HMD with the internal display of the HMD facing the user's eyes or face. In some embodiments, the first display generating component displays a camera view of the physical environment in front of the first user, or includes a transparent or translucent portion of the physical environment in front of the first user that is visible to the first user. In some embodiments, the physical environment made visible to the first user through the first display generating component is a portion of the physical environment in front of the display surface of the second display generating component (e.g., location B7000-b, which includes the area in front of the display surface of the second display generating component and physical object 7014 (and optionally physical object 7028), the area in front of the external display of the HMD, etc.). In some embodiments, the computing system requires that the first display generation component be moved to a predetermined orientation relative to a first user or a specific part of the first user (e.g., the internal display of the HMD is oriented to face the user's eyes or face, the first user moves to face the display surface of the first display generation component, and / or the internal display of the HMD is upright relative to the user's face) in order to trigger the display of a computer-generated experience via the first display generation component.In some embodiments, individual computer-generated experiences are selected according to the current state of the computing system (e.g., one or more states determined based on contextual information (e.g., time, location, what physical objects are present in front of the user, user identification information, new notifications or alerts generated on the computing system, etc.)) when the user begins and / or completes a movement to a pre-defined spatial relationship between the user and the first display generation component, and / or which user interface elements (one or more) (e.g., one or more user interface elements that convey the identification information and characteristics of the selected computer-generated experience, and / or user interface elements that convey the contextual information used for the selected computer-generated experience) are displayed by the second display generation component. In the embodiments shown in Figures 7F and 7G, the current state of the computing system is determined based on the current position of the display generation component(s) and / or physical objects (one or more) present in front of the display surface of the second display generation component (and optionally, the current position and physical objects present in front of the external display of the HMD housing both the first and second display generation components).

[0151] As described above with respect to Figures 7A to 7E and repeatedly stated herein, the first display generation component (e.g., display 7100) and the second display generation component (e.g., display 7102) are shown in Figures 7F to 7I as being located in two separate and isolated parts of the physical environment, but it should be understood that the first and second display generation components are two display generation components that are optionally housed in the same housing (e.g., the housing of a single HMD) or mounted on the same support structure (e.g., back to back or mounted on two faces of a single wall or surface) and facing in different directions (e.g., facing in opposite directions, facing at different angles, etc.). According to some embodiments, the user can move the housing or support structure of the first and second display generation components (e.g., rotate, change direction, or flip vertically or horizontally) to move the first display generation component to the user or to a pre-configured spatial configuration relative to the user or the user's face or eyes. According to some embodiments, the user can move the first display-generating component to a pre-configured spatial configuration relative to the user or the user's face or eyes by inserting the user's head into the housing of the display-generating component or by attaching the support structure of the display-generating component to a part of the user's body (e.g., head, shoulder, nose, ear, etc.). Thus, the simultaneous presence of the first user and the second display-generating component at position B7000-b and the simultaneous presence of the first user and the first display-generating component at position A7000-a represent a first time before the pre-configured spatial relationship between the user and the first display-generating component for triggering the display of a computer-generated experience is satisfied, and the second time when the pre-configured spatial relationship is satisfied by the movement of the user and / or the display-generating component(s), and the available CGR experience is displayed via the first display-generating component.

[0152] In some embodiments, the second display generation component is a low-resolution, smaller, simpler, mono-stereoscopic, monochrome, low-power, and / or secondary display, and the first display generation component is a higher-resolution, larger, more complex, stereoscopic, full-color, full-output, and / or primary display of the computing system. In some embodiments, the second display generation component is used by the computing system to display notifications and prompts to the user to position the first display generation component in a pre-configured spatial relationship to the user's eyes for viewing state information, event information, state information related to the computing system, and in particular additional available content related to the current context. In some embodiments, the second display generation component is used by the computing system when the first display generation component is not positioned in front of the user's eyes (or more generally, when the user is not in a position to fully enjoy the CGR content displayed on the first display generation component), and / or when the computing system's display generation component (e.g., as part of a single HMD) is positioned on a desk, in the user's hand, in a container (e.g., a backpack, holder, case, etc.), or in standby mode (e.g., plugged into a charging station, set to low-power mode, etc.). In some embodiments, while displaying information using the second display generation component, the computing system continues to monitor the spatial relationship between the user (e.g., the first user or any user) and the first display generation component (e.g., using sensors mounted on or surrounded by the housing of the first display generation component (e.g., motion sensors, orientation sensors, image sensors, touch sensors, etc.) and / or external sensors (e.g., motion sensors, orientation sensors, image sensors, etc.)).In some embodiments, upon detecting relative movement between the first display generation component and the user (for example, when the user picks up a display generation component enclosed in the same housing or mounted on the same support structure and turns the display surface of the first display generation component toward the user's eyes or face, and / or when the user puts on an HMD including the first and second display generation components on the user's head), the computing system displays a computer-generated experience corresponding to the state of the computing system at the time the pre-defined spatial relationship between the user and the first display generation component is satisfied (for example, optionally, this is the same state the computing system had at the time the information (e.g., one or more user interface objects indicating the availability of a computer-generated experience) was displayed by the second display generation component at the start of the relative movement).

[0153] In some embodiments, as shown in Figures 7G and 7I, when the computing system is displaying individual computer-generated experiences corresponding to the current context (e.g., the respective states of the computing system as shown by user interface objects 7012 and 7026 in Figures 7F and 7H) via a first display generation component, the computing system optionally uses a second display generation component to display state information (state information 7022 and 7034, respectively). In some embodiments, the displayed state information conveys information about the computer-generated content displayed via the first display generation component, and optionally, the state of the user viewing the computer-generated content via the first display generation component (e.g., the appearance of the user's face or eyes). Other aspects and details relating to the display of state information using the second display generation component while the computing system is displaying computer-generated content using the first display generation component will be described with reference to Figures 7A to 7E and their accompanying descriptions, and the processes described with reference to Figures 8 to 13. In some embodiments, the second display generation component stops displaying any content (e.g., user interface element 7012 or 7026) when the first display generation component begins displaying content and / or when a pre-configured spatial relationship between the user and the first display generation component is met. In some embodiments, the computing system does not display state information or any other content when displaying a computer-generated experience using the first display generation component. In some embodiments, the computing system uses the second display generation component to display other information (e.g., a digital clock, a weather forecast, a countdown timer based on the duration of the computer-generated experience, or time allocated to a first user to use the first display generation component) when displaying a computer-generated experience using the first display generation component.

[0154] In some embodiments, the individual computer-generated experiences displayed via the first display generation component are mixed reality experiences in which virtual content is seen simultaneously with a representation of a physical environment (e.g., location B, a portion of the physical environment in front of the first user). In some embodiments, the representation of the physical environment includes a camera view of the portion of the physical environment that is within the first user's field of view if the user's eyes were not obstructed by the presence of the first and second display generation components (e.g., if the first user was not wearing an HMD or holding an HMD in front of their eyes). In mixed reality mode, the CGR content (e.g., movies, augmented reality environments, user interfaces, and / or virtual objects) is displayed so as to superimpose or replace at least a portion, if not all, of the representation of the physical environment. In some embodiments, the first display generation component includes a transparent portion in which a portion of the physical environment is visible to the first user, and in mixed reality mode, the CGR content (e.g., movies, augmented reality environments, user interfaces, virtual objects, etc.) is projected onto a physical surface or empty space within the physical environment and is seen through the transparent portion together with the physical environment. In some embodiments, the CGR content is displayed on a portion of the display, blocking the view of at least a portion, but not all, of the physical environment visible through the transparent or semi-transparent portion of the first display generating component. In some embodiments, the first display generating component 7100 does not provide a view of the physical environment, but provides a complete virtual environment (e.g., without a camera view or transparent pass-through portion) augmented with a real-time visual representation (one or more) of the physical environment (e.g., a stylized representation or a segmented camera image) as it is currently being captured by one or more sensors (e.g., a camera, a motion sensor, or other attitude sensor).In some embodiments, in mixed reality modes (e.g., augmented reality based on a camera view or transparent display, or augmented virtual reality based on a virtualized representation of a physical environment), the first user is not fully immersed in the computer-generated environment and is still provided with sensory information (e.g., visual, audio, etc.) that directly corresponds to the physical environment surrounding the first user and the first display-generating component. In some embodiments, while the first display-generating component displays a fully immersive environment, the second display-generating component optionally displays state information without information about the user's eye state (e.g., state information only about CGR content), or without displaying any state information at all.

[0155] In some embodiments, the computing system optionally has any number of different states corresponding to the availability of different computer-generated experiences for display via a first display generation component. Each different state of the computing system optionally has a corresponding set of one or more user interface elements displayed by a second display generation component when the computing system enters and / or remains in that state. Each different state of the computing system optionally is triggered by the satisfaction of a corresponding event or set of events and / or a corresponding set of one or more predefined criteria. Although only two states of the computing system, two user interface objects corresponding to the two states, and two different computer-generated experiences are shown in the embodiments described with reference to Figures 7F to 7I, a third state, a third user interface element, and a third computer-generated experience may optionally be implemented by the computing system in a manner similar to that described with respect to the two states, user interface elements, and computer-generated experiences described in the embodiments. In some embodiments, any finite number of states, user interface elements, and computer-generated experiences may optionally be implemented.

[0156] In some embodiments, the computer-generated experience provided by the first display generation component is an immersive experience (e.g., an AR or VR experience) that takes into account the actions of a first user in a physical environment (e.g., gestures, movement, speech, and / or gaze). For example, when the user's hand moves in the physical environment, or when the user moves in the physical environment (e.g., changes direction or walks), the user interface of the computer-generated three-dimensional environment and / or the user's view are updated to reflect the user's hand movement (e.g., pressing and opening a virtual window in the AR environment, activating a user interface element in a home screen or menu presented in the AR environment, etc.) or the user's movement (e.g., the user's viewpoint moves relative to the AR environment or virtual three-dimensional game world, etc.).

[0157] In some embodiments, different computer-generated experiences (e.g., a first computer-generated experience, a second computer-generated experience, etc.) are representations of the same physical environment but are AR experiences that include different virtual elements selected based on the state of the computing system (e.g., indicated by one or more user interface elements displayed by the second display-generated component (e.g., circle 7012, square 7026, etc.)). For example, in some embodiments, the computer-generated experience optionally includes a view of the same room in which a first user is located. Immediately before the user places the display surface of the first display-generated component in front of the user, the computing system determines that it has displayed one of several different event reminders on the second display-generated component, and the computing system displays a representation of the room having one of several different themed virtual wallpapers on the representation of the room's walls, while displaying a separate introductory video for the event corresponding to the individual event reminder.

[0158] In some embodiments, different computer-generated experiences are either augmented reality or virtual reality experiences, depending on the context (e.g., the state of the computing system, which is determined based on relevant contextual information (e.g., location, time, user identification information, receipt of notifications or alerts, etc.) and / or what is shown on a second display generation component). In some embodiments, after a computer-generated experience is initiated in one of the AR and VR modes, the experience can transition to the other of the AR and VR modes (e.g., in response to a user request, in response to the fulfillment of other pre-configured conditions, etc.).

[0159] In some embodiments, the computing system is configured to use a second display generation component to display various user interfaces and / or user interface objects for different applications, based on the state of the computing system. For example, in some embodiments, one or more user interface elements displayed on the second display generation component include elements of an electronic calendar (e.g., a social calendar, a work calendar, a daily planner, a weekly planner, a monthly calendar, a standard calendar showing dates and weeks by month) having scheduled events, appointments, holidays, and / or reminders. In some embodiments, the computing system displays different computer-generated experiences through the first display generation component once a pre-defined spatial configuration between the first display generation component and a first user is met (e.g., the first user or the user's eyes are facing the display surface of the first display generation component, the first user is in a position that allows them to view the content displayed by the first display generation component, etc.), and the particular computer-generated experience displayed is based on what calendar content was displayed on the second display generation component immediately before the movement that placed the first display generation component and the first user in the pre-defined spatial configuration began and / or ended. For example, upon determining that one or more user interface elements displayed on the second display generation component correspond to a first calendar event (e.g., the user interface elements indicate event information, warnings, notifications, calendar data, notes, etc., for the first calendar event), the computing system displays a first computer-generated experience corresponding to the first calendar event (e.g., detailed information and / or interactive information (e.g., preview, video, venue and attendee models, etc.)).Upon determining that one or more user interface elements displayed on the second display generation component correspond to a second calendar event (e.g., the user interface elements indicate event information, warnings, notifications, calendar data, notes, etc., for the second calendar event), the computing system displays a second computer-generated experience corresponding to the second calendar event (e.g., detailed information and / or interactive information (e.g., preview, video, venue and attendee models, etc.)). In some embodiments, when the double-sided HMD is not worn by the user (e.g., the external display is placed on a desk facing the user), the external display of the HMD is used to display a calendar containing the current date, time, weather information, geographical location, and / or a list of tasks or scheduled appointments coming up that day or within a predetermined period (e.g., in the next two hours, in the next five minutes, etc.). When a user picks up the HMD and places the HMD's internal display in front of their eyes (for example, by lifting the HMD or by putting the HMD on the user's head), the HMD's internal display displays calendar details (for example, showing a more complete calendar including the current week or current month, showing all scheduled events for the day, showing more details of upcoming events, etc.). In some embodiments, one or more user interface elements corresponding to a first calendar event include a notification of the first calendar event, and a user interface element corresponding to a second calendar event is a notification of the second calendar event.

[0160] In some embodiments, the computing system uses a second display generation component to display media objects such as photographs and / or video clips having two-dimensional images, and uses a first display generation component to display a three-dimensional experience or un-abbreviated media content corresponding to the media objects displayed on the second display generation component. For example, user interface elements displayed on the second display generation component optionally include snapshots or clips from a long video, reduced-resolution or two-dimensional versions of three-dimensional video, non-interactive user interfaces corresponding to an interactive computer environment, etc., and when criteria for triggering the display of such enhanced content are met (e.g., a pre-configured spatial configuration of the first display generation component and a first user, and optionally, when other conditions are met (e.g., the user is seated, the HMD has sufficient power, etc.)), the first display generation component displays the long video, three-dimensional video, interactive computer environment, etc. In some embodiments, when the double-sided HMD is not worn by the user (e.g., placed on a desk with the external display facing the user), the external display of the HMD is used to display a visual representation of available media items that can be displayed via the internal display of the HMD. In some embodiments, the available media items are modified depending on the current position of the HMD and / or the availability of the media items as specified by the media item provider. When the user picks up the HMD and places the internal display in front of the user's eyes, the first display generation component displays the actual content of the media item (e.g., showing a more complete movie, a more immersive experience, and / or enabling more interactive features of the media item, etc.).

[0161] In some embodiments, the computing system uses a second display generation component to display a warning for an incoming communication request (e.g., an incoming phone call, an audio / video chat request, a video conferencing request, etc.), and the computing system uses the first display generation component to display the corresponding communication environment when the first display generation component is positioned in a pre-configured physical configuration for a first user (e.g., by the movement of the first user, the first display generation component, or both). In some embodiments, the communication environment displayed via the first display generation component represents a simulated environment in which virtual avatars or images of each participant exist (e.g., avatars are seated around a table representation in front of the first user, or are present as speakers on the table surface in front of the first user, etc.). In some embodiments, upon detecting the positioning of the first display generation component in a pre-configured physical configuration for a first user, the computing system accepts the incoming communication request and initiates the corresponding communication session (e.g., using the first display generation component and other components of the computing system). In some embodiments, upon detecting the placement of a first display generation component in a pre-configured physical configuration for a first user, the computing system initiates an application corresponding to an incoming communication request and displays (e.g., using the first display generation component) the user interface of the application, which the first user can choose to accept the incoming communication request. In some embodiments, when the double-sided HMD is not being worn by the user (e.g., placed on a desk with the external display facing the user), the external display of the HMD is used to display a notification of the incoming communication request when such a request is received by the computing system. In some embodiments, the notification provides caller identification information and an indication of the type of communication session being requested.When a user picks up the HMD and places the HMD's internal display in front of their eyes (for example, by lifting the HMD with their hands or by putting the HMD on their head), the HMD's internal display shows a communication interface corresponding to the received communication request, and the user can use the HMD's internal display to begin communicating with the caller. In some embodiments, the computing system starts different applications (or different modes of the same application) depending on the characteristics of the incoming communication request (e.g., caller's identification information, time, subject of the call, etc.). For example, in response to an incoming request from a colleague, the computing system displays a user interface on a first display generation component that awaits pre-configured input from the first user before starting a communication session, and in response to an incoming request from a family member, the computing system starts a communication session without displaying a user interface and / or without requiring pre-configured input from the first user. In another embodiment, in response to an incoming request arriving at the user's home, the computing system starts a communication session with an avatar of the first user in casual attire, and in response to an incoming request arriving at the user's office, the computing system starts a communication session with an avatar of the first user in work attire. In another embodiment, in response to an incoming telephone call request, the computing system displays a detailed speaker representation of each participant, and in response to an incoming video chat request, the computing system displays a full-body representation of each participant showing the actual body movements of the participants. In some embodiments, one or more user interface elements displayed on a second display generation component visually represent specific characteristics of an incoming communication request used by the computing system to determine the characteristics of a computer-generated experience (e.g., the user interface or environment of a communication session).In some embodiments, selected characteristics of the computer-generated experience are also visually represented by one or more user interface elements shown by a second display-generating component before the computer-generated experience is displayed by a first display-generating component. In some embodiments, the computing system modifies the characteristics of the computer-generated experience in accordance with user input received before the computer-generated experience is displayed using the first display-generating component (e.g., touch gestures on the second display-generating component, touch gestures on the housings of the first and / or second display-generating components, air gestures, voice commands, etc.).

[0162] In some embodiments, the computing system modifies the content displayed on the second display-generating component (e.g., one or more user interface elements) in response to various parameters (e.g., user distance, user identification information, user gestures, etc.). For example, upon detecting a first user at a first distance away from the second display-generating component (e.g., the first distance is less than a first threshold distance but greater than a second threshold distance), the computing system displays a first version of one or more user interface elements (e.g., a large, simple icon or text) to indicate the availability of a personalized computer-generated experience; and upon detecting a first user at a second distance away from the second display-generating component (e.g., the second distance is less than a second threshold distance), the computing system displays a second version of one or more user interface elements (e.g., a graphic, more detailed, etc.) to indicate the availability of a personalized computer-generated experience (e.g., when the first user moves closer to the second display-generating component, the display of the first version of one or more user interface elements is replaced). In another embodiment, upon detecting a user within a threshold distance of a second display generation component, the computing system displays a generic version of one or more user interface elements (e.g., a large, simple icon or text) to indicate the availability of a personalized computer-generated experience; and upon detecting user identification information (e.g., upon detecting the user's fingerprints when picking up a first / second display generation component (e.g., an HMD), or upon the user moving closer to the second display generation component), the computing system displays a user-specific version of one or more user interface elements corresponding to the user's identification information (e.g., customized based on user preferences, usage history, demographics, etc.).

[0163] In some embodiments, the computing system displays a user interface containing selectable options (e.g., one or more user interface elements and / or one or more user interface objects other than user interface elements) before detecting that a first display generation component has been placed in a pre-configured physical configuration for a first user, and detects user input in which one or more of the selectable options are selected, the selectable options including preferences for customizing a computer-generated experience corresponding to one or more user interface elements available for display via the first display generation component. Once the first display generation component has been placed in a pre-configured physical configuration for the first user, the computing system displays a computer-generated experience customized based on the user's selected options(s). In some embodiments, the selectable options correspond to a set of two or more modes of the computing system (e.g., AR mode, VR mode, 2D mode, private mode, parental control mode, DND mode, etc.) from which the computer-generated experience can be presented via the first display generation component.

[0164] In some embodiments, one or more user interface elements displayed by the second display generation component include a preview of a three-dimensional experience available for display by the first display generation component. In some embodiments, the preview provided by the second display generation component is a three-dimensional preview that simulates a viewport to the three-dimensional experience. The user can move their head relative to the second display generation component to view different parts of the three-dimensional environment represented by the three-dimensional experience. In some embodiments, the preview is initiated when the user picks up the second display generation component (e.g., picks up a double-sided HMD) and / or places the second display generation component in a pre-configured spatial configuration for the first user (e.g., holds the HMD with the external display facing the user's eyes). In some embodiments, after the preview has been initiated on the first display generation component, the computing system initiates a computer-generated experience on the first display generation component in response to the detection that the user has placed the first display generation component in a pre-configured spatial relationship with the first user (e.g., the user holds the HMD with the internal display facing the user's face or eyes, the user wears the HMD on the user's head, etc.).

[0165] Figures 7H to 7J illustrate how, in some embodiments, different computer-generated experiences are displayed depending on how the first display generation component is maintained in a pre-configured spatial relationship or configuration for a first user during the presentation of a computer-generated experience. In some embodiments, the different computer-generated experiences are related to one another. For example, the different computer-generated experiences may be a preview of a three-dimensional computer-generated experience and the three-dimensional computer-generated experience itself, or a segment or edited version of a computer-generated experience and the full version of a computer-generated experience, respectively. In some embodiments, the computing system determines that the first display generation component is positioned in a pre-configured spatial relationship or configuration for a first user when the display surface of the first display generation component faces the first user or a pre-configured part of the first user (e.g., the user's eyes or face) and / or when the user is in a position to view content displayed on the first display generation component (e.g., the user is wearing an HMD, holding an HMD, lifting an HMD in their hands, placing an HMD on a pre-configured viewing station, connecting an HMD to a pre-configured output device, etc.). In some embodiments, the computing system determines which of the computer-generated experiences to display on the first display generation component (or optionally, the entire HMD surrounding the first display generation component (and optionally a second display generation component)) is being worn by a first user (e.g., secured by straps, remaining in front of the user's eyes without support from the user's hands). In some embodiments, the computing system determines whether the first display generation component is being worn by a first user based on the state of devices or sensors other than the first display generation component (e.g., straps or buckles on the housing of the first display generation component, touch sensors or position sensors attached to the housing of the first display generation component).For example, the strap or buckle optionally has an open state and a closed state, and the strap or buckle is in the closed state when the first display generating component is worn by a first user, and the strap or buckle is in the open state when the first display generating component is only temporarily positioned in front of the user (e.g., held up to eye level by the user's hand) and is not worn by the user. In some embodiments, a touch sensor or position sensor on the housing of the first display generating component switches to a first state ("yes" state) when the housing of the first display generating component is stationary and supported by the user's nose, ears, head and / or other parts of the user's body other than the user's hands, and the touch sensor or position sensor switches to a second state ("no" state) when the housing of the first display generating component is supported by the user's hand(s). In some embodiments, by distinguishing how the first display generation component (or an HMD containing the first display generation component) is maintained in a pre-configured state for a first user in order for the user to view a computer-generated experience displayed on the first display generation component, the computing system can better adjust the interaction model and the depth of the content presented to the user.For example, the computing system enables a first interaction model that requires the user's hand movement as input (e.g., air gestures, microgestures, inputs provided on a control device separate from the housing of the first display generation component) only when the first display generation component is not being maintained by the user's hand(s) in a pre-configured state for the first user. While the first display generation component is being maintained by the user's hand(s) in a pre-configured state for the first user, the computing system enables only other types of interaction models that do not require the user's hand(s) to move away from the housing of the first display generation component (e.g., speech interaction, gaze interaction, touch interaction on the housing of the first display generation component).

[0166] In Figure 7I, the first user 7202 is positioned at position A7000-a alongside the first display generation component (e.g., display 7100, internal display of the HMD, etc.), facing the display surface of the first display generation component. This is to illustrate exemplary scenarios, according to some embodiments, in which the first display generation component is configured in a pre-configured position relative to the first user or a pre-configured portion of the first user (e.g., the first display generation component is positioned in front of the user's face or eyes). In the exemplary scenario shown in Figure 7I, the first user 7202 is holding the sensor object 7016 in the user's hand. This position of the sensor object 7016 relative to the user's hand corresponds to the state of the first display generation component when it is not being worn by the first user 7202. Another exemplary scenario corresponding to the state of the first display generating component when it is not being worn by the first user is when the first display generating component is the HMD's display (e.g., the HMD's internal display, the HMD's single display, etc.) and the HMD's display is held or lifted by the first user's hands so as to face the first user's eyes (as opposed to being supported by the user's head, nose, ears, or other parts of the user that are not the user's hands). In some embodiments, in a scenario where the sensor object 7016 is held in the first user's hands (which also corresponds to the first display generating component being supported by the user's hands or not being worn by the first user when it is positioned in front of the user's face or eyes), the computing system displays a second computer-generated experience 7030 (e.g., an experience corresponding to the second state of the computing system shown in Figure 7H).

[0167] In Figure 7J, the first user is positioned at position A7000-a alongside the first display generating component, facing the display surface of the first display generating component 7100. This is to illustrate another exemplary scenario, according to some embodiments, in which the first display generating component 7100 is in a pre-configured position relative to the first user or a pre-configured part of the first user (for example, the first display generating component is positioned in front of the user's face or eyes). In the scenario shown in Figure 7J, the first user 7202 no longer holds the sensor object 7016 in their hand, but has positioned the sensor object on the user's body (for example, on the user's back), and as a result, the first user 7202 is now wearing the sensor object 7016. This position of the sensor object 7016 relative to the user's hand and body corresponds to the state of the first display generating component 7100 when it is worn by the first user. Another exemplary scenario corresponding to the state of the first display generating component when the first display generating component is worn by a first user is when the first display generating component is the display of the HMD (e.g., the internal display of the HMD) and the HMD is normally worn on the head of the first user (e.g., fully strapped, buckled, stationary to the user's nose, ears, and / or head, as opposed to being supported by the user's hands(s)).In some embodiments, in scenarios where the sensor object 7016 is worn by the user and not held in the first user's hands (also corresponding to the first display generation component being worn by the user and not supported by the user's hands when placed in front of the user's face or eyes), the computing system displays a third computer-generated experience 7036 (which also corresponds to the second state of the computing system shown in Figure 7H (e.g., showing square 7026), but is a different experience from the second computer-generated experience (e.g., experience 7030 in Figure 7I) due to the state of the sensor object 7016 (and accordingly the wearing state of the first display generation component (e.g., display 7100 or internal display of the HMD)).

[0168] As shown in Figure 7J, the computing system detects (e.g., using a camera 7104 and / or other sensors) that a first user 7202 is moving its hand in the air to provide an air gesture, and accordingly moves a virtual object 7032 onto the representation 7028' of the physical object 7028. In some embodiments, the computing system disables at least some of the input devices (e.g., touch-sensitive surfaces, buttons, switches, etc.) provided on the housing of the first display generation component (e.g., display 7100, internal display of the HMD, single display of the HMD, etc.) while displaying a third computer-generated experience 7036. In some embodiments, the computing system enables at least one interaction model (e.g., an interaction model supporting air hand gestures, micro-gestures, and / or inputs detected via input devices separate from the housing of the first display generation component) that was not enabled when the first display generation component was not worn by the first user 7202 (as determined (e.g., based on the state of a sensor object 7016 or other sensors, etc.)). In some embodiments, the third computer-generated experience 7036 and the second computer-generated experience 7030 are related experiences that have corresponding content (e.g., the same content or different versions of the same content) but different interaction models (e.g., different interaction models or overlapping but different sets of interaction models).

[0169] As shown in Figures 7I and 7J, in some embodiments, the computing system optionally includes a second display generation component (e.g., a display housed in a different housing from the first display generation component (e.g., display 7102), or a display housed in the same housing as the first display generation component (e.g., back-to-back or facing in different directions in other ways) (e.g., an HMD with an internal display and an external display)). In some embodiments, the second display generation component optionally displays state information related to the content shown via the first display generation component (e.g., a modified visual representation of the content), and optionally displays state information related to the state of a first user (e.g., an image or representation of the user's face or eyes) and / or the operating mode of the computing system (e.g., mixed reality mode, virtual reality mode, full passthrough mode, parental control mode, private mode, DND mode, etc.). In some embodiments, the second display generation component also displays user interface elements corresponding to computer-generated experiences available for display by the first display generation component. More details of the operation of the second display generation component and the corresponding operation of the first display generation component will be described with reference to the processes described in Figures 7A-7E and 7F-7I, and Figures 8-13. In some embodiments, a computing system that displays different computer-generated experiences based on whether the first display generation component is worn by a user, when the first display generation component is in a pre-configured state for the user, does not have any other display generation components other than the first display generation component and therefore does not display state information and / or user interface elements indicating the availability of the computer-generated experiences described herein.

[0170] In some embodiments, when the first display generating component is positioned in a pre-configured configuration for a first user (e.g., the display surface of the first display generating component faces the user's eyes or face and / or is within a threshold distance from the user's face), the computing system optionally uses the first display generating component to display different types of user interfaces (e.g., system user interfaces (e.g., application launch user interfaces, home screens, multitasking user interfaces, configuration user interfaces, etc.) versus application user interfaces (e.g., camera user interfaces, infrared scanner user interfaces (e.g., showing a heatmap of the current physical environment), augmented reality measurement applications (e.g., automatically displaying measurements of physical objects in the camera view), etc.)) when the first display generating component is positioned in a pre-configured configuration for a first user (e.g., the display surface of the first display generating component faces the user's eyes or face and / or is within a threshold distance from the user's face, etc.), depending on whether the first display generating component is worn by a first user (e.g., the HMD is strapped or buckled to the user's head and can remain in front of the user's eyes without support from the user's hands, or is simply held in front of the user's eyes by the user's hands, etc., and would fall away without support from the user's hands, etc.). In some embodiments, the computing system takes a photograph or video of the physical environment captured in the camera view in response to user input detected via an input device (e.g., a touch sensor, a contact strength sensor, a button, a switch, etc.) located on the housing of the first display generation component, while the computing system is displaying an application user interface using the first display generation component.

[0171] In some embodiments, when determining a response to user input detected while the user is holding the first display-generating component in front of the user and is not wearing the first display-generating component, the computing system prioritizes touch input detected on a touch-based input device located on the housing of the first display-generating component over microgesture input and / or air gesture input detected in front of the first user (e.g., the microgesture input and air gesture input are performed by the user's hands that are not holding the housing of the first display-generating component). In some embodiments, when determining a response to user input detected while the user is wearing the first display-generating component (e.g., when the user's hands do not need to support the first display-generating component), the computing system prioritizes microgesture input and / or air gesture input detected in front of the first user over touch input detected on a touch-based input device located on the housing of the first display-generating component. In some embodiments, upon detecting multiple types of inputs simultaneously (e.g., inputs performed by a hand away from the first display generation component, inputs performed by a hand touching the first display generation component or its housing, etc.), and in accordance with the determination that the first display generation component is being worn by the user while in a pre-configured configuration for the user (e.g., the HMD containing the first display generation component is strapped to the user's head, buckled, not supported by the user's hands, etc.), the computing system enables interaction with the displayed computer-generated experience based on gestures performed by a hand located away from the first display generation component and its housing (e.g., air gestures, micro-gestures, etc.) (e.g., the gestures are captured by a camera on the HMD, a mechanical or touch-sensitive input device, or a sensor worn on the user's hand, etc.).Upon determining that the first display generation component is not being worn by the user while it is in a pre-configured state for the user (e.g., not secured to the user's head with a strap, not buckled, or supported by the user's hands), the computing system enables interaction with the displayed computer-generated experience based on gestures performed by the first display generation component or hands on its housing (e.g., touch gestures, operation of physical controls, etc.) (e.g., gestures are captured by touch-sensitive surfaces on the HMD housing, buttons or switches on the HMD housing, etc.).

[0172] Figures 7K to 7M show a computing system (e.g., computing system 101 in Figure 1, or computing system 140 in Figure 4) that includes at least a first display generation component (e.g., display 7100, internal display of the HMD, single display of the HMD, etc.) and optionally a second display generation component (e.g., display 7102, external display of the HMD, etc.), which displays a computer-generated experience (e.g., augmented reality experience, augmented virtual experience, virtual reality experience, etc.) to the user via the first display generation component (e.g., display 7100, internal display of the HMD, single display of the HMD, etc.) according to physical interactions between the user and physical objects in the physical environment (e.g., picking up a musical instrument and playing it, picking up a book and opening it, holding a box and opening it, etc.). In some embodiments, only specific physical interactions that satisfy pre-set criteria corresponding to physical objects can trigger the display of a computer-generated experience. In some embodiments, different computer-generated experiences are optionally displayed depending on which of a set of criteria is satisfied by the physical interaction with the physical object. In some embodiments, different computer-generated experiences include different augmented reality experiences corresponding to different modes of manipulating physical objects (e.g., tapping, flicking, striking, opening, shaking, etc.). Figures 7K–7M are used to illustrate the processes described later, including the processes shown in Figures 8–13.

[0173] As shown in Figure 7K, the user (e.g., user 7202) is in a scene 105 that includes physical objects in a room with walls and a floor (e.g., objects including a box lid 7042 and box body 7040, books, musical instruments, etc.). In Figure 7K, the user is holding a first display generating component (e.g., a display 7100, HMD, handheld device, etc.) in their hand 7036. In some embodiments, the first display generating component is not held in the user's hand 7036 but is supported by a housing or support structure placed on the user's body (e.g., head, ears, nose, etc.). In some embodiments, the first display generating component (e.g., a head-up display, projector, etc.) is positioned in front of the first user's eyes or face and is supported by another support structure (e.g., a tabletop, TV stand, etc.) that is not part of the user's body.

[0174] In some embodiments, as shown in Figure 7K, the computing system provides an augmented reality view 105' of a physical environment (e.g., a room containing physical objects) via a first display generation component (e.g., a display 7100, a display for an HMD, etc.). In the augmented reality view of the physical environment, the view of a portion of the physical environment includes representations of physical objects (e.g., representations 7042' of the lid 7042 and 7040' of the body 7040 of a box), and optionally, representations of the surrounding environment (e.g., representations 7044' of the support structure 7044 supporting the physical objects, as well as representations of the walls and floor of the room). In addition to representations of physical objects in the environment, the computing system also displays some virtual content (e.g., user interface objects, visual extensions of physical objects, etc.) including visual indications (e.g., labels 7046 or other visual indications) that one or more computer-generated experiences corresponding to physical objects (e.g., a box containing the lid 7042 and body 7040, another physical object in the environment, etc.) are available for display via the first display generation component. As shown in Figure 7K(B), in some embodiments, a visual indication (e.g., label 7046) is displayed in a view of the physical environment at a position corresponding to the location of a physical object (e.g., at the location of the box lid 7042). For example, label 7046 appears to be located on top of the box lid 7042 in a view of the physical environment displayed via a first display generation component.

[0175] In some embodiments, a visual indication (e.g., label 7046, or other visual indication) includes descriptive information about a computer-generated experience (e.g., icons, graphics, text, animations, video clips, images, etc.) that is available for display by the first display-generating component. In some embodiments, when the first display-generating component of the computing system or one or more cameras move within the physical environment, and / or when a physical object moves within the physical environment, the augmented reality view of the physical environment shown by the first display-generating component includes only a representation of less than a threshold portion of the physical object (e.g., less than 50% of the physical object, or without including the main part of the physical object (e.g., box lid 7042, book title text, sound-generating part of a musical instrument, etc.)), and the computing system stops (or stops displaying) the visual indication in the view of the augmented reality environment.

[0176] In some embodiments, the visual indication includes prompt or guidance information regarding physical interactions necessary to trigger the display of a computer-generated experience (e.g., animated figures, indicators pointing to specific parts of a representation of a physical object). In some embodiments, the computing system displays prompt or guidance information regarding physical interactions necessary to trigger the display of a computer-generated experience only in response to detecting certain user inputs that satisfy a first set of criteria (e.g., criteria used to assess whether the user is interested in viewing the computer-generated experience, criteria used to detect the presence of the user, criteria for detecting the user's hand contact on a physical object). As shown in Figure 7L, in some embodiments, when the computing system detects that the user is touching a physical object with their hand but has not performed the interaction necessary to trigger the display of a computer-generated experience (for example, the hand 7038 is touching the lid 7042 or the body 7040 of the physical object, but has not opened the lid 7042), the computing system prompts the user to open the lid 7042 by displaying prompt or guidance information (e.g., an animated arrow 7048, or other visual effects or virtual objects) at a position in the view of the augmented reality environment corresponding to the position of the lid 7042. In some embodiments, the prompt and guidance information (e.g., the direction of the arrow, the sequence of animations, etc.) is updated depending on how the user interacts with the physical object. In some embodiments, a representation of the user's hand (e.g., a representation 7038' of the hand 7038) is shown in the augmented reality view 105' of the physical environment as the user manipulates a physical object in the physical environment using their hand.It should be noted that prompt and guidance information differs from the actual computer-generated experience available for display via the first display generation component when a necessary physical interaction with a physical object is detected (e.g., opening the lid 7042, or any other interaction (e.g., picking up the box 7040 from the support 7044 after removing the lid 7042)). In some embodiments, the computing system does not display any prompt or guidance information in response to a physical interaction with a physical object that does not meet the criteria for triggering the display of the computer-generated experience (e.g., the computing system does not display the computer-generated experience and does not display prompt and guidance information, but optionally maintains the display of a visual indication (e.g., label 7046) to indicate that the computer-generated experience is available for display).

[0177] As shown in Figure 7M, in some embodiments, when the computing system detects that the user has performed a physical interaction with a physical object necessary to trigger the display of a computer-generated experience, the computing system displays the computer-generated experience. For example, in response to detecting that the user's hand 7038 lifts the lid 7042 from the box body 7040, the computing system determines that the necessary physical interaction with the physical object has met a pre-set criterion for triggering the display of the computer-generated experience, and displays the computer-generated experience using a first display generation component (e.g., a display 7100, an internal display of the HMD, a single display of the HMD, etc.). In Figure 7M, the computer-generated experience is an augmented reality experience that displays representations 7038' of the user's hand 7038, 7042' of the box lid 7042, and 7040' of the box body 7040 in positions and orientations corresponding to their physical positions and orientations in the physical environment. In addition, in some embodiments, the augmented reality experience also displays virtual content (e.g., a virtual ball 7050 that appears to pop out of the box body 7040 and casts a virtual shadow inside the box body, a virtual platform 7052 that replaces the representation of the physical support 7044 beneath the box 7040, etc.) simultaneously with the representation of the physical environment. In addition, in some embodiments, the representation of walls in the physical environment is replaced with a virtual overlay of the augmented reality experience. In some embodiments, once the computer-generated experience is initiated, as the user continues to interact with physical objects, the computing system displays changes in the augmented reality environment according to the user's physical manipulation of physical objects (e.g., the box body 7040, the box lid 7042, etc.) and optionally according to other inputs detected through various input devices of the computing system (e.g., gesture input, touch input, gaze input, voice input, etc.). In some embodiments, the computer-generated experience progresses in a manner determined according to the continued physical interaction with physical objects.For example, in response to detecting that the user moves the box lid 7042 in the physical environment, the computing system moves the representation 7042' of the box lid 7042, pushing the virtual ball 7050 into the empty space above the representation 7040' of the box body 7040. In response to detecting that the user places the box lid 7042 back on the box body 7040, the computing system displays the representation 7042' of the box lid 7042 back on the representation 7040' of the box body 7040, and stops displaying the virtual ball 7050. In some embodiments, the computing system requires the user to be in physical contact with a physical object when performing the necessary physical interaction to trigger the display of the computer-generated experience, and the computing system stops displaying the computer-generated experience upon determining that the user has ceased physical contact with the physical object for longer than a threshold time. For example, in some embodiments, the computing system stops displaying the computer-generated experience immediately after detecting that the physical object has been released from the user's hands. In some embodiments, the computing system stops displaying the computer-generated experience when it detects that a physical object has landed on another physical surface and come to rest after being released from the user's hands.

[0178] In some embodiments, visual feedback provided in response to the detection of a user's physical interaction with a physical object prior to a criterion for triggering the display of a computer-generated experience includes a preview of the computer-generated experience and has visual characteristics that are dynamically updated according to the characteristics of the physical interaction once the physical interaction is detected. For example, animation, visual effects, and / or the extent of the virtual object (e.g., size, dimensions, angular range, etc.), the amount of detail in the visual feedback, and the brightness, color saturation, visual sharpness, etc. of the visual feedback are optionally adjusted according to characteristic values ​​of the interaction with the physical object in the physical environment (e.g., characteristic values ​​include distance traveled, angular range of travel, speed of travel, type of interaction, distance to a given reference point, etc.) (e.g., dynamically in real time, periodically, etc.). For example, in some embodiments, if the physical object is a book, as the cover of the book is slowly opened by the user in the physical environment, the colors and light of the computer-generated experience emerge from the gap between the cover and the first page, becoming brighter and more saturated as the cover is opened further. A complete computer-generated experience is optionally initiated in a three-dimensional environment when the book cover is opened beyond a threshold amount and a first criterion is met. In another embodiment, when the user lifts the corner of the box lid 7042, virtual light is shown to emerge from the representation 7040' of the box body 7040, and a virtual ball 7050 is shown to be briefly visible. As the user lifts the corner of the box lid 7042 higher, more virtual light is shown to emerge from the representation 7040' of the box body 7040, and the virtual ball 7050 begins to move around within the representation 7040' of the box body 7040. When the user finally lifts the box lid 7042 away from the box body 7040, the computer-generated experience is initiated, the entire three-dimensional environment changes, the representation of the room is replaced with a virtual platform 7052, and the virtual ball 7050 flies out of the box representation 7040'.

[0179] In some embodiments, a computer-generated experience is optionally triggered by two or more types of physical interaction. In other words, the criteria for triggering a computer-generated experience associated with a physical object are optionally met by a first method of interacting with the physical object and a second method of interacting with the physical object. For example, a computer-generated experience associated with a book is optionally initiated when the user picks up the book and places it on a bookshelf with the cover upright in front of the user's face, and when the user picks up the book and opens the cover with their hand. In some embodiments, a computer-generated experience is optionally initiated from a different part of the computer-generated experience. For example, the criteria for triggering a computer-generated experience associated with a physical object are optionally met by the same method of interacting with the physical object, but with different parameter values ​​(e.g., different pages, different speeds, different times, etc.). For example, a computer-generated experience associated with a book may optionally begin with a first portion of the experience depending on whether the user picks up the book and opens it from the first page, or optionally begin with a second, different portion depending on whether the user picks up the book and opens it from a previously bookmarked page. In another embodiment, slowly opening the book may trigger a computer-generated experience with calm background music and / or more subdued colors, while quickly opening the book may trigger a computer-generated experience with more lively background music and brighter colors. The book embodiment is merely illustrative. The same principle applies to other computer-generated experiences and other triggering physical interactions associated with other types of physical objects. In some embodiments, different computer-generated experiences may be associated with the same physical object and triggered by different ways of interacting with the physical object.For example, a box could be associated with two different computer-generated experiences, the first of which is triggered when the user opens the lid of the box (e.g., a virtual ball pops out of the box when the user presses the lid), and the second of which is triggered when the user inverts the box (e.g., a virtual insect emerges from the bottom of the box and follows the user's finger as it moves across the bottom of the box). In some embodiments, different ways of interacting with a physical object trigger different versions of the computer-generated experience, enabling different input modalities. For example, if a book is held in one hand and opened by the other, one-handed air gestures (e.g., air tap gestures, waving, sign language gestures, etc.) and microgestures are enabled to interact with the computer-generated experience, while touch gestures are not enabled to interact with the computer-generated experience. When the book is held open with both hands, air gestures are disabled, and touch gestures on the back, front, and / or sides of the book (e.g., tap, swipe, etc.) are enabled to interact with the computer-generated experience.

[0180] Figures 7N to 7Q illustrate how, in some embodiments, the system chooses to perform or not perform an action depending on the input detected on the housing of the display generation component, depending on whether one or both hands are detected on the housing at the time of input. In some embodiments, touch input as an input modality is disabled if both hands are detected simultaneously on the housing of the display generation component (e.g., on the housing of the HMD housing the display generation component, on the frame of the display generation component, etc.). In some embodiments, the computing system will respond to touch input performed by either of the user's hands on the housing of the display generation component (e.g., on the structure of the display generation component or on touch-sensing surfaces or other touch sensors located around it) as long as only one hand (e.g., the hand providing the touch input) is touching the housing of the display generation component. If the computing system detects that an additional hand is also touching the housing of the display generation component at the time the touch input is detected, the computing system will ignore the touch input and refrain from performing the action corresponding to the touch input (e.g., activating controls, interacting with the computer-generated experience, etc.). In some embodiments, the computing system detects the presence of both hands using the same sensors and input devices used to detect touch input provided by one hand. In some embodiments, a camera is used to capture the position of the user's hands, and the images from the camera are used by a computing system to determine whether both of the user's hands are touching the housing of the display generating component when touch input is detected by one or more touch sensors present on the housing of the display generating component. In some embodiments, other means are employed to detect touch input and / or whether one or both of the user's hands are touching or supporting the housing of the display generating component.For example, according to various embodiments, the presence and orientation of a user's hand(s) on the housing of the display generation component can be detected using position sensors, proximity sensors, mechanical sensors, etc.

[0181] In Figure 7N, the user (e.g., user 7202) is present in a physical scene 105 that includes physical objects (e.g., a box 7052), walls 7054 and 7056, and a floor 7058. The user 7202 is standing in front of a display generation component 7100 (e.g., a tablet display device, a projector, an HMD, an internal display of an HMD, a head-up display, etc.) supported by a stand. The display generation component 7100 displays a view of the physical environment 105. For example, in some embodiments, the view of the physical environment is a camera view of the physical environment. In some embodiments, the view of the physical environment is provided through a transparent portion of the display generation component. In some embodiments, the display generation component 7100 displays an augmented reality environment having both a view of the physical environment and virtual objects and content displayed superimposed on or replacing a portion of the view of the physical environment. As shown in Figure 7N, the view of the physical environment includes a representation 7052' of the box 7052, representations 7054' and 7056' of the walls 7054 and 7056, and a representation 7058' of the floor 7058. In Figure 7N, the user 7202 is not touching the housing of the display generation component, and the computing system does not detect any touch input on the housing of the display generation component. In some embodiments, a representation 7038' of the user's hand 7038 is shown via the display generation component 7100 as part of the representation of the physical environment. In some embodiments, the display generation component 7100 displays the virtual environment or displays nothing before any touch input is detected on its housing.

[0182] In some embodiments, in addition to touch input, the computing system is optionally configured to detect hover input near the housing of the display generation component. In some embodiments, proximity sensors positioned on the housing of the display generation component are configured to detect a user's finger or hand approaching the housing of the display generation component and to generate an input signal based on the proximity of the finger or hand to the housing of the display generation component (e.g., proximity to a part of the housing configured to detect touch input, to other parts of the housing, etc.). In some embodiments, the computing system is configured to detect hover input at different locations near the housing of the display generation component (e.g., using proximity sensors located in different parts of the housing of the display generation component) and to provide different feedback according to the location of the hover input. In some embodiments, the computing system adjusts the values ​​of various characteristics of the visual feedback based on the hover distance of the detected hover input (e.g., the distance of the fingertip(s) from the surface of the housing or the touch-sensitive part of the housing).

[0183] As shown in Figure 7O, in some embodiments, the computing system detects that the hand 7038 or a portion thereof (e.g., the fingers of the hand 7038, or two fingers of the hand 7038) has moved within a threshold distance from the first touch-sensitive portion of the housing of the display generation component 7100 (e.g., the upper left edge portion of the housing of the display generation component 7100, or the left or upper edge portion of the HMD). In response to the detection that the hand 7038 or a portion thereof has moved within a threshold distance from the first touch-sensitive portion of the housing of the display generation component 7100, the computing system optionally displays one or more user interface objects at a position corresponding to the position of the hand 7038 or a portion thereof. For example, the user interface object 7060 is displayed near the upper left edge portion of the housing of the display generation component 7100 (or the left or upper edge portion of the HMD) next to the hand 7038 or the raised fingers of the hand 7038. In some embodiments, the computing system dynamically updates the position of one or more displayed user interface objects in accordance with changes in the position of the hand 7038 or a portion thereof near the housing of the display generation component (e.g., moving closer to / further away from it, and / or moving up and down along the left edge, etc.) (e.g., moving user interface object 7060 closer to or further away from the left edge of the housing, and / or moving the left edge of the housing up and down). In some embodiments, the computing system dynamically updates the appearance of one or more displayed user interface objects in accordance with changes in the position of the hand or a portion thereof near the housing of the display generation component (e.g., changing the size, shape, color, opacity, resolution, content, etc.).In some embodiments, the computing system optionally changes the type of user interface object(s) displayed near the first touch-sensitive portion of the housing of the display-generating component in accordance with changes in the posture of the hand 7038 near the housing of the display-generating component (e.g., lifting two fingers instead of one, lifting different fingers, etc.) (e.g., changing from displaying a representation of a first control (e.g., volume control, power control, etc.) to displaying a representation of a second control (e.g., display brightness control, WiFi control, etc.), changing from displaying a first type of affordance (e.g., scroll bar, button, etc.) to displaying a second type of affordance (e.g., scroll wheel, switch, etc.)). In some embodiments, one or more user interface objects evolve from mere indications of a control (e.g., small dots, faint shadows, etc.) to more specific and distinct representations of a control (e.g., buttons or switches with graphics and / or text on them indicating the state of the computing system or the display-generating component) as the hand 7038 moves closer to the first touch-sensitive portion of the housing of the display-generating component. In some embodiments, one or more user interface objects (e.g., user interface object 7060, or other user interface objects) are displayed simultaneously with the view of the physical environment. In some embodiments, one or more user interface objects are displayed simultaneously with the virtual environment if the virtual environment was displayed before, for example, the hover input by hand 7038 was detected. In some embodiments, one or more user interface objects are displayed without simultaneous display of the view of the physical environment or any other virtual content if, for example, the display generation component was not displaying either the view of the physical environment or any other virtual content before the hover input was detected.

[0184] In some embodiments, as shown in Figure 70, the computing system detects a touch input on the first touch-sensitive portion on the housing of the display-generating component, after detecting a hover input near the first touch-sensitive portion on the housing of the display-generating component, for example. In some embodiments, upon detection of a touch input on the first portion of the touch-sensitive portion of the housing of the display-generating component, and without another hand being detected on the housing of the display-generating component, the computing system performs a first action corresponding to the touch input. For example, as shown in Figure 70, upon detection of a touch input on the upper portion of the left edge of the housing of the display-generating component, the computing system activates an augmented reality experience that includes some virtual content (e.g., a virtual overlay 7062) in combination with a view of the physical environment (e.g., the virtual overlay 7062 is displayed in a position corresponding to the position of the box 7052, and as a result, the overlay 7062 appears to be positioned in front of the representation 7052' of the box 7052 in the augmented reality environment). In some embodiments, an initial touchdown of the user's hand 7038 or a portion of the hand only triggers visual feedback that a touch has been detected on the housing of the display generating component, and the touch has not yet met the criteria for initiating any particular action on the computing system (e.g., user interface objects (e.g., user interface object 7060) appear in their fully functional state if they were not displayed or were simply displayed as an indication, and user interface objects (e.g., user interface object 7060) change their appearance to indicate that a touch has been detected, etc.).In some embodiments, the computing system evaluates the touch input provided by the hand 7038 against criteria for detecting one or more valid touch inputs, and, according to the determination that the touch input satisfies criteria for triggering a first action associated with one or more user interface objects (e.g., user interface object 7060 or another user interface object displayed in place of user interface object 7060) (e.g., lowering display brightness, lowering volume, switching to AR mode, etc.), the computing system performs the first action associated with one or more user interface objects, and, according to the determination that the touch input satisfies criteria for triggering a second action associated with one or more user interface objects (e.g., user interface object 7060 or another user interface object displayed in place of user interface object 7060) (e.g., increasing display brightness, increasing volume, switching to VR mode, etc.), the computing system performs the second action associated with one or more user interface objects.

[0185] In some embodiments, as shown in Figure 7P, the computing system detects a user's hand 7038 approaching a second touch-sensitive portion of the housing of the display generation component (e.g., the lower left edge of the housing of the display generation component, the lower left edge of the HMD, the upper right edge of the HMD, etc.) and optionally displays one or more user interface objects (e.g., user interface object 7064 or other user interface objects) accordingly. When the computing system detects a touch input on the second touch-sensitive portion of the housing of the display generation component, and based on the determination that only one hand (e.g., hand 7038) is touching the housing of the display generation component at the time the touch input is detected, the computing system performs a second action corresponding to the touch input performed by hand 7038 or a portion thereof. In this embodiment, activation of control 7064 initiates an augmented reality experience that includes some virtual content (e.g., a virtual ball 7066, other virtual objects, etc.) displayed in combination with a view of the physical environment (for example, the virtual ball 7066 is displayed in a position corresponding to the position of the floor 7058, and as a result, the virtual ball 7066 appears to be placed on the surface of the representation 7058' of the floor 7058). Other features related to this exemplary scenario are similar to those described with respect to Figure 70, the only difference being that the hand 7038 is detected near and / or on the first touch-sensitive portion of the housing of the display-generating component, and certain visual feedback and actions(s) performed correspond to touch input detected on or near the first touch-sensitive portion of the housing of the display-generating component. For brevity, similar features will not be repeated herein.

[0186] Figure 7Q shows that when touch input from hand 7038 is detected on the housing of the display generation component, another hand (e.g., hand 7036) is touching (e.g., holding, supporting, or otherwise touching) the housing of the display generation component (e.g., display 7100, HMD, etc.). In Figure 7Q, hand 7038 is touching the first touch-sensitive portion and the second touch-sensitive portion of the housing of the display generation component, respectively, providing the same touch input that previously triggered the execution of the first action (e.g., displaying the virtual overlay 7062 in Figure 7O) and the second action (e.g., displaying the virtual ball 7066 in Figure 7P), respectively. However, once one of the touch inputs from hand 7038 is detected, while another hand (e.g., hand 7036) is also touching the housing of the display generation component, the computing system, according to the determination that a touch input has been detected (e.g., the determination that the other hand is touching the opposite side of the housing from the side where the touch input was detected, or the determination that the other hand is touching somewhere on the housing, etc.), stops performing the action corresponding to the detected touch input. For example, as shown in Figure 7Q, the computing system still optionally displays a user interface object (e.g., user interface object 7060, user interface object 7064, etc.) near the location of the hand 7038 that provided a valid touch input, but does not initiate the corresponding computer-generated experience when the other hand 7036 is touching the housing of the display generation component. In some embodiments, the computing system does not display any user interface object if hover inputs from two hands are detected simultaneously near the touch-sensitive portion of the housing of the display generation component, and / or if touch inputs are detected simultaneously on two sides of the housing of the display generation component.In some embodiments, if two hands are detected simultaneously approaching and / or touching a display generation component (e.g., an HMD, other types of display generation components), the user is more likely to want to remove the display generation component rather than provide input to it. Therefore, it is more advantageous to ignore such hover or touch input without explicit user instruction (e.g., to reduce user confusion, save power, etc.).

[0187] In some embodiments, actions performed in accordance with touch input detected on the housing of a display generation component change the state of the computing system. For example, according to a determination that the touch input meets a first criterion, the computing system switches to a first state, and according to a determination that the touch input meets a second criterion different from the first criterion, the computing system switches to a second state different from the first state. In some embodiments, the first and second criteria have different position-based criteria, requiring that the touch input be detected at different locations on the housing. In some embodiments, the first and second criteria have different intensity-based criteria, requiring that the touch input meet different intensity thresholds. In some embodiments, the first and second criteria have different duration-based criteria, requiring that the touch input be detected less than a threshold displacement over different threshold times on the housing. In some embodiments, the first and second criteria have different distance-based criteria, requiring that the touch input move more than a different threshold distance. In some embodiments, the first and second criteria have different touch pattern criteria, requiring that t...

Claims

1. A computing system comprising a first display generation component, a second display generation component, and one or more input devices, wherein the first display generation component and the second display generation component are enclosed in the same housing. The computer-generated environment is displayed via the first display generation component, wherein the first display generation component displays hardware components of the computing system that are visible to the user of the computing system while the user of the computing system is located in an individual physical environment. Displaying the computer-generated environment via the first display generation component while displaying state information corresponding to the computing system via the second display generation component, wherein the second display generation component is a hardware component of the computing system that is visible to the individual person while the individual person other than the user of the computing system is present with the user in the individual physical environment, and the display of the state information is A visual representation of a portion of the user of the computing system who is positioned to view the computer-generated environment via the first display generation component, It provides visual indication of content within the computer-generated environment and includes one or more graphic elements that are different from the visual representation of a portion of the user. This includes displaying them simultaneously, Detecting individual events, In response to detecting the aforementioned individual event, To change the immersion level of the computer-generated environment displayed via the first display generation component, This includes changing the appearance of the visual representation of a portion of the user of the computing system, and changing the state information displayed via the second display generation component, Methods that include...

2. The computing system is configured to display the computer-generated environment at at least a first immersion level, a second immersion level, and a third immersion level. The method according to claim 1.

3. In response to detecting the individual events, the immersion level of the computer-generated environment displayed via the first display generation component may be changed. In accordance with the determination that the individual event satisfies the first criterion, the computer-generated environment is switched from being displayed at the second immersion level to being displayed at the first immersion level. In accordance with the determination that the individual event is an event that satisfies a second criterion different from the first criterion, the computer-generated environment is switched from being displayed at the second immersion level to being displayed at the third immersion level. The method according to claim 2, including the method described in claim 2.

4. Changing the immersion level of the computer-generated environment displayed via the first display generation component includes switching from displaying the computer-generated environment at a second immersion level to displaying the computer-generated environment at a first immersion level, wherein the computer-generated environment displayed at the first immersion level provides a view of the individual physical environment having computer-generated content below a threshold amount. Changing the state information displayed via the second display generation component includes switching from displaying the computer-generated environment at a second immersion level to displaying the computer-generated environment at a first immersion level, and switching from displaying the visual representation of a portion of the user of the computing system together with one or more graphic elements that provide the visual indication of the content in the computer-generated environment, to displaying the visual representation of a portion of the user without the one or more graphic elements. The method according to claim 1.

5. The method according to claim 1, wherein the computing system simultaneously displays one or more graphic elements that provide a visual indication of the visual representation of the portion of the user and the content in the computer-generated environment, based on the determination that the computer-generated environment is a mixed reality environment including representations of the individual physical environments surrounding the first display generation component and at least a threshold amount of virtual objects.

6. Changing the immersion level of the computer-generated environment displayed via the first display generation component includes switching from displaying the computer-generated environment at a second immersion level to displaying the computer-generated environment at a third immersion level, wherein the computer-generated environment displayed at the third immersion level provides a virtual environment having a representation less than a threshold amount of the individual physical environment. Changing the state information displayed via the second display generation component includes switching from displaying the computer-generated environment at a second immersion level to displaying the computer-generated environment at a third immersion level, and switching from displaying the visual representation of a portion of the user of the computing system together with one or more graphic elements that provide the visual indication of the content in the computer-generated environment, to displaying the one or more graphic elements without the visual representation of the portion of the user. The method according to claim 1.

7. The first user request to activate the privacy mode of the computing system is detected, which requires that one or more graphic elements displayed via the second display generation component have visibility below a first threshold visibility, while the computer-generated environment is displayed via the first display generation component and the state information corresponding to the computing system is displayed via the second display generation component. In response to detecting the first user request, While the visibility of one or more graphic elements providing visual indication of content in the computer generation environment exceeds the first threshold visibility corresponding to the privacy mode, the visibility of one or more graphic elements on the second display generation component is reduced to below the first threshold visibility corresponding to the privacy mode, in accordance with the determination that a first user request has been received. If the visibility of one or more graphic elements that provide visual indication of content in the computer-generated environment does not exceed the first threshold visibility corresponding to the privacy mode, then, in accordance with the determination that a first user request has been received, the visibility of one or more graphic elements is maintained below the first threshold visibility corresponding to the privacy mode. The method according to claim 1, including the method described in claim 1.

8. The privacy mode requires that the visual representation of the portion of the user displayed via the second display generation component has a visibility level below a second threshold visibility level, and the method In response to detecting the first user request, In accordance with the determination that a first user request has been received while the visibility of the visual representation of the portion of the user exceeds the second threshold visibility corresponding to the privacy mode, the visibility of the visual representation of the portion of the user on the second display generation component is reduced to below the second threshold visibility corresponding to the privacy mode. In accordance with the determination that a first user request has been received while the visibility of the visual representation of the portion of the user does not exceed the second threshold visibility corresponding to the privacy mode, the visibility of the visual representation of the portion of the user is maintained below the second threshold visibility corresponding to the privacy mode. The method according to claim 7, including the method described in claim 7.

9. While the privacy mode is active on the computing system, the system detects a second separate event that alters the level of immersion of the computer-generated environment displayed via the first display generation component, In response to detecting the second individual event and the corresponding change in the immersion level of the computer-generated environment displayed via the first display generation component, the state information displayed via the second display generation component is to be stopped from being modified. The method according to claim 7, including the method described in claim 7.

10. While the computer-generated environment is being displayed via the first display generation component, a second user request is detected to activate the silent (DND) mode of the computing system. In response to detecting the second user request, a visual indicator is displayed via the second display generation component to indicate that the silent mode is active, The method according to claim 1, including the method described in claim 1.

11. While the silent mode is active on the computing system, the system detects a third separate event that alters the level of immersion of the computer-generated environment displayed via the first display generation component. In response to detecting the third individual event and the corresponding change in the immersion level of the computer-generated environment displayed via the first display generation component, the state information displayed via the second display generation component is stopped from being modified. The method according to claim 10, including the method described in claim 10.

12. While the computer-generated environment is being displayed via the first display generation component, a third user request is detected to activate a parental control mode of the computing system, which requires that one or more graphic elements displayed via the second display generation component have a visibility greater than a third threshold visibility. In response to detecting the third user request, While the visibility of one or more graphic elements providing visual indication of content in the computer generation environment is below the third threshold visibility corresponding to the parental control mode, according to the determination that a third user request has been received, the visibility of one or more graphic elements on the second display generation component is increased to exceed the third threshold visibility corresponding to the parental control mode. The method according to claim 1, including the method described in claim 1.

13. While the parental control mode is active on the computing system, the system detects a fourth distinct event that alters the level of immersion of the computer-generated environment displayed via the first display generation component, In response to detecting the fourth individual event and the corresponding change in the immersion level of the computer-generated environment displayed via the first display generation component, the state information displayed via the second display generation component is to be stopped from being modified. The method according to claim 12, including the method described in claim 12.

14. The method according to claim 1, wherein the simultaneous display of the visual representation of the portion of the user and the one or more graphic elements includes displaying the visual representation of the portion of the user at a first depth and displaying the one or more graphic elements at a second depth smaller than the first depth from an external viewpoint of the state information displayed via the second display generation component.

15. The method according to claim 1, wherein one or more graphic elements providing a visual indication of the content in the computer-generated environment include at least a progress indicator showing the progress of the content in the computer-generated environment when displayed via the first display generation component.

16. The method according to claim 1, wherein one or more graphic elements that provide a visual indication of content in the computer-generated environment have a first display characteristic, the value of which is based on the value of the first display characteristic of the content in the computer-generated environment.

17. The method according to claim 1, wherein the one or more graphic elements that provide a visual indication of the content in the computer-generated environment include one or more sub-parts of the content in the computer-generated environment.

18. The method according to claim 1, wherein one or more graphic elements that provide a visual indication of the content in the computer-generated environment include metadata that identifies the content in the computer-generated environment.

19. Displaying one or more graphic elements that provide visual indication of content within the computer-generated environment, To detect the first movement of the individual person who is at a position to view the state information displayed via the second display generation component to the second display generation component, In response to detecting the first movement of the individual person relative to the second display generation component, and determining that the distance between the individual person and the second display generation component has decreased from above the first threshold distance to below the first threshold distance, the display of one or more graphic elements is updated to increase the information density of the visual indication of the content in the computer-generated environment provided by the one or more graphic elements. The method according to claim 1, including the method described in claim 1.

20. The computer-generated environment is displayed via the first display generation component, and the state information corresponding to the computing system is displayed via the second display generation component, while a fifth individual event is detected that is triggered by the individual person who is at a position to view the state information displayed via the second display generation component, In response to detecting the fifth individual event, and in accordance with the determination that the fifth individual event satisfies the fourth criterion, which requires that, as a result of the fifth individual event, a pre-defined measure of interaction increases from below a pre-defined threshold to above a pre-defined threshold, and the computer-generated environment is displayed at a third immersion level, The immersion level of the computer-generated environment displayed via the first display generation component is changed from the third immersion level to the second immersion level, wherein the computer-generated environment displayed at the second immersion level includes a representation of an increased amount of the individual physical environment compared to the computer-generated environment displayed at the third immersion level. The method according to claim 1, including the method described in claim 1.

21. In response to the detection of the fifth individual event, and in accordance with the determination that the fifth individual event satisfies the fourth criterion, Modifying the state information displayed via the second display generation component, which includes increasing the visibility of the visual representation of the portion of the user of the computing system in conjunction with changing the immersion level of the computer-generated environment displayed via the first display generation component from the third immersion level to the second immersion level, The method according to claim 20, including the method described in claim 20.

22. The method according to claim 21, wherein detecting the fifth individual event includes detecting the individual person entering a predetermined area surrounding the user of the computing system.

23. The method according to claim 21, wherein detecting the fifth individual event includes detecting that the individual person performs a predetermined gesture toward the user of the computing system.

24. One or more processors, The first display generation component, A second display generation component, which is enclosed in the same housing as the first display generation component, A computing system comprising a memory that stores one or more programs configured to be executed by one or more processors, wherein the one or more programs are The computer-generated environment is displayed via the first display generation component, wherein the first display generation component displays hardware components of the computing system that are visible to the user of the computing system while the user of the computing system is located in an individual physical environment. Displaying the computer-generated environment via the first display generation component while displaying state information corresponding to the computing system via the second display generation component, wherein the second display generation component is a hardware component of the computing system that is visible to the individual person while the individual person other than the user of the computing system is present with the user in the individual physical environment, and the display of the state information is A visual representation of a portion of the user of the computing system who is positioned to view the computer-generated environment via the first display generation component, It provides visual indication of content within the computer-generated environment and includes one or more graphic elements that are different from the visual representation of a portion of the user. This includes displaying them simultaneously, Detecting individual events, In response to detecting the aforementioned individual event, To change the immersion level of the computer-generated environment displayed via the first display generation component, This includes changing the appearance of the visual representation of a portion of the user of the computing system, and changing the state information displayed via the second display generation component, A computing system that includes instructions to perform a task.

25. The computing system according to claim 24, wherein one or more programs include instructions for performing the method described in any one of claims 2 to 23.

26. A first display generation component, a second display generation component enclosed in the same housing as the first display generation component, and a computing system capable of communicating with one or more input devices, when executed by the computing system, The computer-generated environment is displayed via the first display generation component, wherein the first display generation component displays hardware components of the computing system that are visible to the user of the computing system while the user of the computing system is located in an individual physical environment. Displaying the computer-generated environment via the first display generation component while displaying state information corresponding to the computing system via the second display generation component, wherein the second display generation component is a hardware component of the computing system that is visible to the individual person while the individual person other than the user of the computing system is present with the user in the individual physical environment, and the display of the state information is A visual representation of a portion of the user of the computing system who is positioned to view the computer-generated environment via the first display generation component, It provides visual indication of content within the computer-generated environment and includes one or more graphic elements that are different from the visual representation of a portion of the user. This includes displaying them simultaneously, Detecting individual events, In response to detecting the aforementioned individual event, To change the immersion level of the computer-generated environment displayed via the first display generation component, This includes changing the appearance of the visual representation of a portion of the user of the computing system, and changing the state information displayed via the second display generation component, A program that includes instructions to perform an action.

27. The program according to claim 26, which, when executed by the computing system, further includes an instruction causing the computing system to perform the method according to any one of claims 2 to 23.

Citation Information

Patent Citations

  • Communication terminal, interview system, display method, and program

    JP2017069827A

  • Controls and interfaces for user interaction in virtual spaces

    JP2019536131A

  • Controls and Interfaces for User Interactions in Virtual Spaces

    US20180095636A1