Head-mounted display with pass-through image processing

The head-mounted display system addresses the issue of obstructed real-world view in VR by using stereo imaging and corrected depth maps to seamlessly integrate VR and real-world environments, enhancing user interaction and navigation.

JP2026027228APending Publication Date: 2026-02-18VALVE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025169824
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-14
Filing Date
2025-10-08
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Head-mounted displays (HMDs) used in virtual reality (VR) environments obstruct the user's view of the real world, causing difficulties in navigating and interacting with the physical environment, and existing pass-through image processing technologies exhibit rough response times and distortion, leading to disorientation and discomfort.

Method used

A head-mounted display system that includes cameras to capture real-world images, processes depth information using stereo imaging, and projects corrected images onto a depth map or 3D mesh, allowing seamless integration of VR and real-world environments through trigger-activated pass-through modes.

Benefits of technology

Enables users to interact with and navigate the real world while immersed in VR without removing the HMD, maintaining immersion and reducing disorientation by providing accurate, distortion-free real-world views.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026027228000001_ABST
    Figure 2026027228000001_ABST
Patent Text Reader

Abstract

To provide a head-mounted display (HMD) for use in a virtual reality (VR) environment in which a user can view the real world without removing the head-mounted display.SOLUTION: The HMD104 includes a display 108 for providing virtual content and / or images to the user 100 and image capture devices, such as the first camera 110 and / or the second camera 112. The first camera 110 and / or the second camera 112 are disposed in or near the front of the MD104 to capture images of the surroundings 102 and pass through the images of the surroundings to the user for viewing on the display. The environment 102 includes a computer 150 that communicatively couples to the HMD, the controller 106, the tracking system 122, and / or the remote computing resources 142 via a network 124 and / or wired technologies.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This is a PCT application claiming priority to U.S. Patent Application No. 16 / 848,332, filed April 14, 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 837,668, filed April 23, 2019. Applications Nos. 16 / 848,332 and 62 / 837,668 are incorporated herein by reference in their entireties. [Technical Field]

[0002] [Background technology] Head-mounted displays are used in a variety of applications, including engineering, medicine, the military, and video games. In some cases, head-mounted displays may present information or images to a user as part of a virtual reality or augmented reality environment. For example, a user may wear a head-mounted display while playing a video game to immerse the user in a virtual environment. Head-mounted displays provide an immersive experience but block views of the physical or real world. As a result, users may find it difficult to pick up objects (e.g., a controller) and / or recognize other individuals in the real world. Additionally, users may not be aware of physical boundaries in the real world (e.g., walls). While removing the head-mounted display may allow a user to see, constantly taking it off and putting it back on may be cumbersome, require the user to reorient themselves between the virtual environment and the real world, and / or otherwise distract from the virtual reality experience. [Brief explanation of the drawings]

[0003] The detailed description will now be described with reference to the accompanying drawings, in which the leftmost digit(s) of a reference number identifies the figure in which the reference number first appears. The same or similar reference numbers in different figures indicate similar or identical items.

[0004] [Figure 1] 1 illustrates a user wearing an exemplary head-mounted display in an exemplary environment, according to an exemplary embodiment of the present disclosure.

[0005] [Figure 2] FIG. 2 is a perspective view of the head-mounted display of FIG. 1 according to one embodiment of the present disclosure.

[0006] [Figure 3] 2 illustrates a user wearing the head-mounted display of FIG. 1 and the head-mounted display presenting a pass-through image, according to one embodiment of the present disclosure.

[0007] [Figure 4] 2 illustrates a user wearing the head-mounted display of FIG. 1 and the head-mounted display presenting a pass-through image, according to one embodiment of the present disclosure.

[0008] [Figure 5] 1. An exemplary process for presenting a pass-through image using the head-mounted display of FIG. 1 will be described according to one embodiment of the present disclosure.

[0009] [Figure 6] 1. An exemplary process for presenting a pass-through image using the head-mounted display of FIG. 1 will be described according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] A head-mounted display is worn by a user to view and interact with content in a virtual reality environment. To provide an immersive experience, the head-mounted display may cover most or even all of the user's field of view. Therefore, the head-mounted display may block the user's vision from the real world, which may result in the user tripping over objects, bumping into furniture, or being unable to recognize individuals in the real world. By removing the head-mounted display, the user can see the real world, but the user must reorient themselves between the virtual reality environment and the real world. Some head-mounted displays may include, for example, an opening visor to allow the user to peer into the real world. However, such a solution disrupts the user's immersion in the virtual reality environment.

[0011] In an attempt to overcome these deficiencies, some head-mounted displays may enable pass-through image processing, allowing each user to view the real world without removing their respective head-mounted display. However, existing pass-through image processing tends to exhibit fairly rough response times and may fail to depict the real world from the user's perspective or point of view and / or may be distorted. As a result, the user may lose balance, become dizzy, disoriented, or become ill.

[0012] This application describes, in part, a head-mounted display (HMD) for use in a virtual reality (VR) environment. The systems and methods described herein may determine information about the real-world environment surrounding a user, the user's location within the real-world environment, and / or the user's posture within the real-world environment. Such information may enable the HMD to display images of the real-world environment in a pass-through manner without distracting the user from the VR environment. In some cases, the HMD may pass-through images of the real-world environment based on one or more trigger events. For example, while wearing the HMD and immersed in the VR environment, the user may activate a button that allows the user to look around the real-world environment. As an additional example, if the user hears something suspicious or interesting, if the user wishes to locate an item (e.g., a controller) within the real-world environment, and / or if a visitor enters the user's vicinity, the HMD may display content associated with the real-world environment. In some cases, the content may be presented to the user in an aesthetic manner to limit distraction from the VR environment. For example, the content may be provided as an overlay to virtual content associated with the VR environment. In such an example, the user can remain immersed in the VR environment while continuing to wear the HMD. Thus, the HMD according to the present application can enhance the user experience when transferring or displaying content between a real-world environment and a VR environment.

[0013] The HMD may include a front end with a display worn on the face adjacent to the user's eyes. The display may output images (or other content) for the user to view. As an example, a user may wear the HMD to play games or view media content (e.g., movies).

[0014] The HMD may include a camera that captures images of the real-world environment. In some cases, the camera may be attached to the display and / or integrated into the front of the HMD. Additionally, or alternatively, the camera may be forward-facing so as to capture images outside the HMD and in front of the user. Furthermore, in some cases, the camera may be separate from the HMD and located throughout the environment or elsewhere on the user (e.g., at the waist).

[0015] In some cases, the cameras may be spatially separated such that the optical axes of the cameras are parallel and separated by a known distance. Thus, the cameras may capture images of the real-world environment from slightly different viewing points. The diversity of information between the viewing points can be used to calculate depth information of the real-world environment (i.e., stereo camera image processing). For example, the HMD and / or a communicatively coupled computing device (e.g., a game console, a personal computer, etc.) can use the image data captured by the cameras to generate depth information associated with the real-world environment.

[0016] For example, the cameras may include a first camera and a second camera horizontally and / or vertically displaced from each other above the front of the HMD. In some cases, the first camera may be located above the front on a first side of the HMD, and the second camera may be located above the front on a second side of the HMD. However, as noted above, the first camera and the second camera may be located elsewhere in the environment and / or above other portions of the user.

[0017] The image data captured by the first camera and / or the second camera may represent different views of a real-world environment (e.g., a room). By comparing the images (or image data) captured by the first camera and / or the second camera, the HMD (and / or another communicatively coupled computing device) can determine the difference or disparity (e.g., using a disparity mapping algorithm). The disparity may represent the difference in coordinates of corresponding image points in the two images. Because the disparity (or disparity value) is inversely proportional to depth in the real-world environment, the HMD and / or another communicatively coupled computing device, such as a game console, may determine depth information associated with the real-world environment (or a portion thereof). In some cases, the depth information may be from the user's perspective (i.e., the user's line of sight).

[0018] Using the depth information, the HMD and / or another communicatively coupled computing device, such as a game console, may generate a depth map or three-dimensional (3D) mesh of the real-world environment (or a portion thereof). For example, the depth map may represent the distance between the user and objects in the real-world environment (e.g., walls of a room, furniture, etc.). Additionally or alternatively, the HMD may include other sensors utilized to generate the depth map and / or 3D mesh. For example, the HMD may include a depth sensor for determining the distance between the user and objects in the real-world environment and may determine that the user is in proximity (e.g., a predetermined proximity or threshold proximity) to an object or boundary of the real-world environment. However, the HMD or game console may additionally or alternatively use light detection and ranging (LIDAR), ultrasonic ranging, stereoscopic ranging, structured light analysis, dot projection, particle projection, time-of-flight observations, etc. for use in generating the depth map and / or 3D mesh.

[0019] Upon generating a depth map and / or 3D mesh of the real-world environment, the HMD and / or another communicatively coupled computing device, such as a game console, may project image data onto the depth map and / or 3D mesh. In this sense, the image data may first be utilized to generate the depth map and / or 3D mesh and may subsequently be superimposed, overlaid, or projected onto the depth map and / or 3D mesh. By doing so, the HMD may display content to the user in a manner that depicts the real-world environment.

[0020] In some cases, the depth map, the 3D mesh, and / or the images captured by the first camera and / or the second camera may be modified to take into account the user's pose or viewpoint. For example, the camera may not be aligned (e.g., horizontally, vertically, and depthwise) with the user's eyes (i.e., the camera is not at the exact position of the user's eyes), so the depth map and / or the 3D mesh may take this discrepancy into account. In other words, the image data captured by the first camera and the second camera may represent a perspective different from the user's perspective. Failure to take this discrepancy into account may describe an incomplete real-world environment, and the user may find it difficult to pick up an object because the depth values ​​or image data are not from the user's viewpoint. For example, because the camera captures image data from a perspective and / or depth different from that of the user's eyes (i.e., the camera is not in the same horizontal, vertical, and / or depth position as the user's eyes), the 3D mesh can take this offset into account to accurately depict and present an undistorted view of the real-world environment to the user. That is, the depth map, 3D mesh, and / or image data may be corrected based at least in part on the difference (or offset) in the coordinate positions (or locations) of the first and second cameras and the user's eyes (e.g., the first and / or second eyes) or field of view. Thus, the image of the real-world environment displayed to the user may accurately represent objects in the real-world environment from the user's perspective.

[0021] Additionally or alternatively, in some cases, the HMD and / or the real-world environment may include sensors that track the user's line of sight, gaze, and / or field of view. For example, the HMD may include an interpupillary distance (IPD) sensor to measure the distance between the pupils of the user's eyes and / or other sensors that detect the direction of the user's gaze. Such sensors can be utilized to determine the user's gaze point so as to accurately portray an image of the real-world environment to the user.

[0022] Pass-through image processing may allow a user to interact with and view objects in a real-world environment, such as a coworker, a computer screen, a mobile device, etc. In some cases, a user may be able to toggle or switch between a VR environment and a real-world environment, or the HMD may automatically switch between the VR environment and the real-world environment in response to one or more trigger events. That is, an HMD may include one or more modes, such as a pass-through mode in which a real-world environment is presented to the user and / or a virtual reality mode in which a VR environment is presented to the user. As an example, while wearing an HMD, a user may desire to drink water. Rather than removing the HMD, the user may actuate (e.g., double-press) a button on the HMD, which causes the real-world environment to be displayed. The display may then present an image of the real-world environment, allowing the user to locate their own glass of water, all without removing the HMD. Then, after locating the glass of water, the user may actuate (e.g., single-press) a button to cause the display to present virtual content of the VR environment. Thus, a pass-through image representing a real-world environment may enable a user to move around the real-world environment and locate objects without bumping into them (e.g., furniture). As a further example, the pass-through image may represent another individual entering the user's real-world environment. Here, the HMD may detect the other individual and present an image on the display such that the user can recognize or be made aware of the other individual.

[0023] In some cases, content associated with the real-world environment displayed to the user may be partially transparent to maintain the user's sense of being in the VR environment. In some cases, content associated with the real-world environment may be combined with virtual content, or only content associated with the real-world environment may be presented (e.g., 100 percent pass-through image processing). Additionally or alternatively, in some cases, images captured by a camera may be presented over the entire display or within a specific portion. Furthermore, in some cases, content associated with the real-world environment may be displayed with dotted lines to indicate to the user which content is part of the real-world environment and which content is part of the VR environment. Such presentation may allow the user to see approaching individuals, disturbances, and / or objects surrounding the user. Regardless of the particular implementation or configuration, the HMD may function to display content associated with the real-world environment to alert, detect, or otherwise recognize objects coming into the user's field of view.

[0024] In some cases, the HMD and / or other computing devices associated with the VR environment, such as a game console, may operate in conjunction with a tracking system in the real-world environment. The tracking system may include sensors that track the user's position within the real-world environment. Such tracking may be used to determine information about the real-world environment surrounding the user, such as the user's location within the environment and / or the user's posture or viewpoint, while the user is immersed in the VR environment. Within the real-world environment, the tracking system may determine the user's location and / or posture. In some cases, the tracking system may determine the user's location and / or posture relative to the center of the real-world environment.

[0025] In some cases, the tracking system may include lighting elements that emit light (e.g., visible or invisible) into the real-world environment and sensors that detect incident light. In some cases, to detect the user's location and / or posture, the HMD may include markers. When light is projected into the real-world environment, the markers may reflect the light, and the sensors may capture the incident light reflected by the markers. The captured incident light can be used to track and / or determine the location of the markers in the environment, which can be used to determine the user's location and / or posture.

[0026] In some cases, the user's location and / or posture within the real-world environment can be utilized to present warnings, instructions, or content to the user. For example, if a user approaches a wall in the real-world environment and knows the user's and the wall's location (via a tracking system), the HMD may display an image representing the wall within the real-world environment. That is, in addition to or instead of determining that the user is approaching a wall using images captured by the camera (i.e., via depth values), the tracking system can determine the user's relative location within the real-world environment. Such tracking can assist in presenting images that accurately correspond to the depth (or placement) of objects within the real-world environment.

[0027] Additionally, images or data obtained from the tracking system can be used to generate a 3D model (or mesh) of the real-world environment. For example, by knowing the location and / or pose of the user, HMD, tracking system, game console, and / or another communicatively coupled computing device, the user's relative location within the real-world environment can be determined. This user's location and / or pose can be used to determine the corresponding portion of the 3D model of the real-world environment that the user is viewing (i.e., the user's field of view). For example, when a camera captures an image that is not associated with the user's viewpoint, the tracking system may determine the user's gaze, pose, or viewpoint. Such information can be used to determine where the user is looking within the real-world environment. Knowing where the user is looking within the real-world environment can be used to correct the images captured by the camera. In doing so, the HMD may accurately display the real-world environment.

[0028] The HMD, tracking system, game console, and / or another communicatively coupled computing device may also compare the depth map and / or 3D mesh generated using the camera image data to the 3D model to determine the user's relative location within the real-world environment. Regardless of the particular implementation, knowing the user's location and / or pose within the real-world environment, the HMD and / or another communicatively coupled computing device may convert points of the depth map and / or points of the 3D mesh into a 3D model of the real-world environment. Images captured by the camera may then be projected onto the 3D model corresponding to the user's viewpoint.

[0029] Therefore, in light of the above, the present application contemplates an HMD that provides pass-through image processing to enhance the VR experience. Pass-through image processing can provide a relatively seamless experience when displaying content from a VR environment and content associated with a real-world environment. Such pass-through image processing provides a less intrusive and unobtrusive solution for viewing content associated with the real-world environment. In some cases, information or content associated with the real-world environment can be selectively provided to the user in response to a trigger, including, but not limited to, a movement, a sound, a gesture, a preconfigured event, a change in the user's movement, etc. Furthermore, in some cases, an HMD according to the present application can take many forms, including a helmet, a visor, goggles, a mask, glasses, and other head or eyewear worn on the user's head.

[0030] The present disclosure provides a general understanding of the principles of the structures, functions, devices, and systems disclosed herein. One or more examples of the present disclosure are illustrated in the accompanying drawings. Those skilled in the art will understand that the devices and / or systems specifically described herein and illustrated in the accompanying drawings are non-limiting embodiments. Features illustrated or described in connection with one embodiment, including between systems and methods, can be combined with features of other embodiments. Such modifications and variations are intended to be within the scope of the appended claims.

[0031] 1 illustrates a user 100 residing in an environment 102 and wearing an HMD 104. In some cases, the user 100 can wear the HMD 104 to immerse the user 100 in the VR environment. In some cases, the user 100 can interact in the VR environment using one or more controllers 106. The HMD 104 includes a display 108 for providing virtual content and / or images to the user 100 and, in some cases, an image capture device, such as a first camera 110 and / or a second camera 112.

[0032] The first camera 110 and / or the second camera 112 may capture images of the environment 102 and pass the images of the environment 102 through to the user 100 for viewing on the display 108. That is, as discussed in detail herein, the images captured by the first camera 110 and / or the second camera 112 may be presented to the user 100 in a pass-through manner to enable the user 100 to view the environment 102 without having to leave the VR environment and / or remove the HMD 104.

[0033] In some cases, the first camera 110 and / or the second camera 112 may be disposed in or near the front of the HMD 104. In some cases, the first camera 110 and / or the second camera 112 may represent a stereo camera, an infrared (IR) camera, a depth camera, and / or any combination thereof. Images captured by the first camera 110 and / or the second camera 112 may represent the environment 102 surrounding the user 100. In some cases, the first camera 110 and / or the second camera 112 may face forward to capture images of the environment 102 in front of the user 100. In some cases, the first camera 110 and / or the second camera 112 may be spatially separated such that the optical axes of the cameras are parallel. Thus, images captured by first camera 110 and / or second camera 112 may represent environment 102 from different viewing points and may be used to determine depth information associated with environment 102 (i.e., stereo camera image processing). However, in some cases, first camera 110 and / or second camera 112 may be located elsewhere within environment 102. For example, first camera 110 and / or second camera 112 may be located on the floor of environment 102, on a desk within environment 102, or the like.

[0034] As described, the HMD 104 may include processor(s) 114 that perform or otherwise implement operations associated with the HMD 104. For example, the processor(s) 114 may cause the first camera 110 and / or the second camera 112 to capture images, subsequently receive the images captured by the first camera 110 and / or the second camera 112, compare the images (or image data), and determine differences between the images (or image data). If the differences are inversely proportional to the depth of objects in the environment 102 (e.g., walls, furniture, a TV, etc.), the processor(s) 114 may determine depth information associated with the environment 102.

[0035] In some cases, using the depth information, the processor(s) 114 may generate a depth map or a 3D mesh of the environment 102. For example, as described, the HMD 104 includes a memory 116 that stores or otherwise has access to the depth map 118 of the environment 102 and / or the 3D mesh 120 of the environment 102. If the image data captured by the first camera 110 and / or the second camera 112 represents a portion of the environment 102, then the depth map 118 and / or the 3D mesh 120 may represent the portion of the environment 102 accordingly. In some cases, upon generating the depth map 118 and / or the 3D mesh 120, the processor(s) 114 may store the depth map 118 and / or the 3D mesh 120 in the memory 116.

[0036] If the image data, depth map, and / or 3D mesh captured by the first camera 110 and / or the second camera 112 are not from the perspective of the user 100 (i.e., the point of view of the user 100), the HMD 104 (and / or another communicatively coupled computing device) may take into account the placement of the first camera 110 and / or the second camera 112 relative to the perspective of the user 100 (e.g., relative to the first eye and / or the second eye, respectively). That is, whether located on the HMD 104 or elsewhere in the environment, the first camera 110 and / or the second camera 112 do not capture images corresponding to the point of view of the user 100 and / or the point of view of the user 100. Thus, the image data, or points in the depth map and / or the 3D mesh may be modified or offset to account for this displacement.

[0037] In some cases, the HMD 104 may operate in association with a tracking system 122. In some cases, the HMD 104 may be communicatively coupled to the tracking system 122 via a network 124. For example, the HMD 104 and the tracking system 122 may include one or more interfaces, such as a network interface 126 and / or a network interface 128, respectively, to facilitate wireless connection to the network 124. The network 124 represents any type of communication network, including data and / or voice networks, and may be implemented using wired infrastructures (e.g., cable, CAT5, fiber optic cable, etc.), wireless infrastructures (e.g., RF, cellular, microwave, satellite, Bluetooth, etc.), or the like. TM ), and / or other connection technologies.

[0038] Tracking system 122 may include components that determine or track the pose of user 100, HMD 104, first camera 110, and / or second camera 112 within environment 102. In this sense, tracking system 122 can determine the location, orientation, and / or pose of user 100, HMD 104, first camera 110, and / or second camera 112 at the time that first camera 110 and / or second camera 112 captured an image of environment 102 for pass-through to user 100. For example, tracking system 122 (and / or another computing device) can analyze and parse images captured by tracking system 122 to identify user 100 and / or the pose of user 100 within environment 102. For example, in some cases, tracking system 122 may include projector(s) 130 and / or sensor(s) 132 that operate to determine the location, orientation, and / or posture of user 100. As shown, in some cases, tracking system 122 may be mounted to a wall of environment 102. Additionally or alternatively, tracking system 122 may be mounted elsewhere within environment 102 (e.g., a ceiling, a floor, etc.).

[0039] Projector(s) 130 are configured to generate and project light and / or images into environment 102. In some cases, the images may include images of visible light perceptible to user 100, images of visible light imperceptible to user 100, images with non-visible light, or combinations thereof. Projector(s) 130 may be implemented with any number of technologies capable of generating and projecting images into / onto environment 102. Suitable technologies include digital micromirror device (DMD), liquid crystal on silicon display (LCOS), liquid crystal display, 3LCD, etc.

[0040] The sensor(s) 132 may include high-resolution cameras, infrared (IR) detectors, sensors, 3D cameras, IR cameras, RGB cameras, etc. The sensor(s) 132 are configured to image the environment 102 in visible light wavelengths, non-visible light wavelengths, or both. The sensor(s) 132 may be configured to capture information to detect the depth, location, orientation, and / or posture of objects within the environment 102. For example, as the user 100 maneuvers around the environment 102, the sensor(s) 132 may detect the location, orientation, and / or posture of the user 100. In some cases, the sensor(s) 132 may capture some or all angles and locations within the environment 102. Alternatively, the sensor(s) 132 may focus or capture images within a defined area of ​​the environment 102.

[0041] The projector(s) 130 and / or the sensor(s) 132 may operate in conjunction with the marker(s) 134 of the HMD 104. For example, the tracking system 122 may project light onto the environment 102 via the projector(s) 130, and the sensor(s) 132 may capture reflected images of the marker(s) 134. Using the captured images, the tracking system 122, such as the processor(s) 136 of the tracking system 122, may determine distance information to the marker(s) 134. Additionally or alternatively, the tracking system 122 may detect the posture (e.g., orientation) of the user 100 within the environment 102. In some cases, the marker(s) 134 may be used to determine the viewpoint of the user 100. For example, the distance between the marker(s) 134 and the eyes of the user 100 may be known. In capturing image data of the marker(s) 134, the tracking system 122 (and / or other communicatively coupled computing devices) may determine the relative viewpoint of the user 100. Thus, the tracking system 122 can utilize the marker(s) 134 of the HMD 104 to determine the relative location and / or posture of the user 100 within the environment 102.

[0042] To define or determine characteristics about the environment 102, upon initiating the gaming application, the HMD 104 may request the user 100 to define a boundary, perimeter, or area of ​​the environment 102 that the user 100 can navigate within while immersed in the VR environment. As one example, the processor(s) 114 may cause the display 108 to present instructions to the user 100 to walk around the environment 102 and define the boundary of the environment 102 (or the area that the user 100 will navigate within while immersed in the VR environment). As the user 100 walks around the environment 102, the HMD 104 may capture images of the environment 102 via the first camera 110 and / or the second camera 112, and the tracking system 122 may track the user 100. In that regard, upon determining the boundary of the environment 102, the tracking system 122 may determine the location of the center of the area (e.g., an origin). Knowing the location of the center of the area may allow the HMD 104 to properly display the relative location of an object or scene within the environment 102. In some cases, the center location may be represented as (0,0,0) in an (X,Y,Z) Cartesian coordinate system.

[0043] In some cases, tracking system 122 may transmit the boundary and / or center location to HMD 104. For example, processor(s) 114 of HMD 104 may store the boundary and / or center origin in memory 116 as indicated by boundary 138. Additionally or alternatively, in some cases, images captured by first camera 110 and / or second camera 112 may be associated with images captured by tracking system 122. For example, images captured by first camera 110 and / or second camera 112 may be used to determine the depth of environment 102 while defining a zone. A depth map may be generated. Accordingly, these depth maps may be associated with specific locations / poses within the environment 102, as determined by tracking the user 100 throughout the environment 102. In some cases, these depth maps may be combined or otherwise used to generate a 3D model or mesh of the environment 102. Upon receiving subsequent image data from the HMD 104 and / or the tracking system 122, the location of the user 100 within the environment 102 may be determined, which may be useful for determining depth information within the environment 102 and from the user's perspective, location, or pose. For example, if the image data captured by the first camera 110 and / or the second camera 112 does not correspond to the user's 100's perspective, the tracking system 122 may determine the user's 100's perspective via images captured from the marker(s) 134. This perspective may be used to modify the image data captured by the first camera 110 and / or the second camera 112 to represent the user's 100's perspective.

[0044] For example, as the user 100 engages with the VR environment and maneuvers around the environment 102, the tracking system 122 may determine the relative location of the HMD 104 within the environment 102 by comparing reflected light from the marker(s) 134 to a central location. Using this information, the HMD 104, the tracking system 122, and / or another communicatively coupled computing device (e.g., a game console) can determine the distance of the HMD 104 from the central location or the location of the HMD 104 relative to the central location. Additionally, the HMD 104, the tracking system 122, and / or another communicatively coupled computing device may determine the pose, such as the viewpoint, of the user 100 within the environment 102. For example, the tracking system 122 may determine that the user 100 is looking toward the ceiling, walls, or floor of the environment 102. Such information can be used to correct or otherwise take into account the position of the first camera 110 and / or the second camera 112 relative to the user's eyes or viewpoint.

[0045] In some cases, tracking system 122 may be coupled to a chassis that has a fixed orientation, or the chassis may be coupled to actuator(s) 140 such that the chassis may move. Actuators 140 may include piezoelectric actuators, motors, linear actuators, and other devices configured to displace or move the chassis or components of tracking system 122, such as projector(s) 130 and / or sensor(s) 132.

[0046] The HMD 104 may additionally or alternatively operate in conjunction with a remote computing resource 142. The tracking system 122 may also be communicatively coupled to the remote computing resource 142. In some examples, the HMD 104 and / or the tracking system 122 may be communicatively coupled to the remote computing resource 142, given that the remote computing resource 142 may have computer processing capabilities that far exceed those of the HMD 104 and / or the tracking system 122. Thus, the HMD 104 and / or the tracking system 122 may utilize the remote computing resource 142 to perform relatively complex analyses of the environment 102 and / or to generate a model (or mesh) of image data of the environment 102. For example, the first camera 110 and / or the second camera 112 may capture image data that the HMD 104 may provide to the remote computing resource 142 via the network 124 for analysis and processing. In some cases, the remote computing resource 142 may transmit content to the HMD 104 for display. For example, in response to a trigger event (e.g., a button press) that configures the HMD 104 to pass through images captured by the first camera 110 and / or the second camera 112, the remote computing resource 142 may transmit a depth map, a 3D mesh, and / or a 3D model onto which the HMD 104 projects the images captured by the first camera 110 and / or the second camera 112. Thus, the images captured by the first camera 110 and / or the second camera 112 for presentation to the user 100 may be transmitted to the remote computing resource 142 for processing, and the HMD 104 may receive content to be displayed on the display 108.

[0047] As described, the remote computing resource 142 includes a processor(s) 144 and a memory 146, which may store or otherwise have access to some or all of the components described with respect to the memory 116 of the HMD 104. For example, the memory 146 may have access to and utilize the depth map 118, the 3D mesh 120, and / or the boundary 138. The remote computing resource 142 may additionally or alternatively store or otherwise have access to the memory 148 of the tracking system 122.

[0048] In some cases, remote computing resources 142 may be remote from environment 102, and HMD 104 may be communicatively coupled to them via network interface 126 and over network 124. Remote computing resources 142 may be implemented as one or more servers and, in some cases, may form part of a network-accessible computing platform implemented as a computing infrastructure of processors, storage, software, data access, etc., maintained and accessible over a network such as the Internet. Remote computing resources 142 do not require end-user knowledge of the physical location and configuration of the systems delivering the services. Common terms associated with these remote computing resources 142 may include “on-demand computing,” “software as a service (SaaS),” “platform computing,” “network-accessible platform,” “cloud services,” “data center,” etc.

[0049] In some cases, environment 102 may include computer 150 (or a gaming application, game console, gaming system) communicatively coupled to HMD 104, controller 106, tracking system 122, and / or remote computing resource 142 via network 124 and / or wired technology. In some cases, computer 150 may perform some or all of the processes described herein, such as those performable by HMD 104, controller 106, tracking system 122, and / or remote computing resource 142. For example, computer 150 may modify image data received from, or depth maps associated with, first camera 110 and / or second camera 112 to take into account the viewpoint of user 100. In such an example, computer 150 may determine by or receive from tracking system 122 indications regarding the viewpoint of user 100 and / or the origin of an area in which the user is located (e.g., the origin of the real-world environment) and / or the origin of the virtual world. Thus, the computer may modify the image data received from, or the depth map associated with, the first camera 110 and / or the second camera 112 based on instructions received from the tracking system 122. In some embodiments, the time(s) (e.g., timestamp(s)) at which the image data was captured by one or both of the cameras 110, 112 may be used to determine the time difference between the time of capturing the image data and the time of displaying the captured image data. For example, the image data that is ultimately displayed to the user may be delayed by tens of milliseconds after the image data is captured by the camera(s) 110 and / or 112.Adjustments can be made to the image data to account for this temporal parallax or mismatch, such as by modifying the pixel data (e.g., through rotational and / or translational adjustments) to present images as they appear in the physical world at the time the images are displayed on the HMD, thereby avoiding the presentation of images that appear to lag behind the user's head movements. In some embodiments, camera(s) 110 and / or 112 can be tracked and timestamp information maintained so that the pose of the camera(s) at the time the image(s) are captured is known, and so that the pose of the camera(s) at the time the resulting image data is received and processed can be used to accurately represent the images spatially, instead of relying on camera movement. Additionally, computer 150 may store or otherwise have access to some or all of the components described with respect to memories 116, 146, and / or 148.

[0050] As used herein, a processor, such as processor(s) 114, 136, 144 and / or a computer processor, may include multiple processors and / or a processor with multiple cores. Additionally, a processor(s) may include one or more cores of different types. For example, a processor(s) may include an application processor unit, a graphics processing unit, etc. In one implementation, a processor(s) may include a microcontroller and / or a microprocessor. A processor(s) may include a graphics processing unit (GPU), a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively, or in addition, functionally described herein may be implemented at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc. In addition, each of the processor(s) may have its own local memory which may also store program components, program data, and / or one or more operating systems.

[0051] As used herein, memory, such as memories 116, 146, 148 and / or the memory of computer 150, may include volatile and nonvolatile memory, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program components, or other data. Such memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, RAID storage systems, or any other medium that can be used to store desired information and that can be accessed by a computing device. Memory may be implemented as a computer-readable storage medium (“CRSM”), which may be any available physical medium accessible by a processor(s) to execute instructions stored in the memory. In one basic implementation, the CRSM may include random access memory (“RAM”) and flash memory. In other implementations, the CRSM may include, but is not limited to, read-only memory ("ROM"), electrically erasable programmable read-only memory ("EEPROM"), or any other tangible medium that can be used to store desired information and that can be accessed by a processor(s).

[0052] 2 illustrates a perspective view of the HMD 104. The HMD 104 may include a front 200 and a rear 202 secured to the head of the user 100. For example, the HMD 104 may include a strand, cord, section, strap, band, or other member operatively connecting the front 200 of the HMD 104 and the rear 202 of the HMD 104. The front 200 includes a display 108 positioned in front of or above the eyes of the user 100 to render images output by an application (e.g., a video game). As discussed above, the display 108 outputs images (frames) viewed by the user 100, allowing the user 100 to perceive the images as if immersed in a VR environment.

[0053] In some cases, the front 200 of the HMD 104 may include a first camera 110 and / or a second camera 112. The first camera 110 and / or the second camera 112 may capture images outside the HMD 104 (i.e., of the environment 102) for the user 100 to view on the display 108. As described above, the optical axes of the first camera 110 and the second camera 112 may be parallel and separated by a predetermined distance so that images captured by the first camera 110 and / or the second camera 112 are used to generate depth information of at least a portion of the environment 102. However, in some cases, as discussed above, the HMD 104 may not include the first camera 110 and / or the second camera 112. For example, the first camera 110 and / or the second camera 112 may be located elsewhere within the environment 102.

[0054] In either scenario, the first camera 110 and / or the second camera 112 capture images that may not correspond to the viewpoint of the user 100 (i.e., when the first camera 110 and / or the second camera 112 are not at the actual position of the user's 100 eyes). However, using the techniques described herein, the image data captured by the first camera 110 and / or the second camera 112 can be modified to take into account the displacement of the first camera 110 and / or the second camera 112 relative to the viewpoint of the user 100.

[0055] For example, to determine the viewpoint of user 100, HMD 104 may include marker(s) 134. As shown in FIG. 2 , in some cases, marker(s) 134 may include a first marker 204(1), a second marker 204(2), a third marker 204(3), and / or a fourth marker 204(4) disposed at the corners, edges, or along the perimeter of front 200. However, marker(s) 134 may be located elsewhere on HMD 104, such as along the top, sides, or rear 202. In some cases, marker(s) 134 may include infrared elements, reflectors, digital watermarks, and / or images that respond to electromagnetic radiation (e.g., infrared) emitted by projector(s) 130 of tracking system 122. Additionally or alternatively, marker(s) 134 may include tracking beacons that emit electromagnetic radiation (e.g., infrared light) that is captured by sensor(s) 132 of tracking system 122. That is, projector(s) 130 may project light into environment 102, and marker(s) 134 may reflect the light. Sensor(s) 132 may capture the incident light that is reflected by marker(s) 134 and tracking system 122, or another communicatively coupled computing device, such as remote computing resource 142, may track and plot the location of marker(s) 134 within environment 102 to determine the movement, position, posture, and / or orientation of user 100. Thus, the marker(s) 134 can be used to indicate the viewpoint of the user 100 for use in correcting, adjusting, or otherwise adapting image data from the first camera 110 and / or the second camera 112 before it is displayed to the user 100.

[0056] Figure 3 illustrates a user 100 wearing an HMD 104 within an environment 300. Figure 3 illustrates the display 108 of the HMD 104, which displays virtual content 302 and interacts within the VR environment while the user 100 is wearing the HMD 104. As discussed above, the HMD 104 may display the virtual content 302 as the user 100 moves around the environment 102.

[0057] The tracking system 122 can be positioned within the environment 300 to track the user 100 from one location to another and determine the user's 100's viewpoint within the environment. For example, the tracking system 122 can utilize the marker(s) 134 on the HMD 104 to determine the user's 100's pose (e.g., location and orientation). In some cases, the user's 100's pose may be relative to a central location of the environment 300. For example, the central location may have coordinates (0,0,0), and using reflected light from the marker(s) 134, the tracking system 122 (or the remote computing resource 142) can determine the user's 100's pose within coordinate space (e.g., (X,Y,Z)). The tracking system 122 can also determine the depth of objects within the environment 300 from the user's 100's perspective or viewpoint. Additionally or alternatively, image data received from the first camera 110 and / or the second camera 112 can be used to determine depth within the environment 300.

[0058] The HMD 104 may adjust the display of the virtual content 302 with objects (e.g., furniture, walls, etc.) within the environment 300 and / or within and / or in front of the user 100. In other words, the objects of the environment 300 (or real world) displayed to the user 100 may correspond to the user's 100's viewpoint within the environment 300. For example, the HMD 104 may display objects within the environment 300 using image data (e.g., pass-through images) captured by the first camera 110 and / or the second camera 112 based at least in part on one or more trigger events. That is, the HMD 104 may be configured to display images on the display 108 at locations that correspond to the actual placement of the objects within the environment 300. Utilizing a depth map generated from image data from the first camera 110 and / or the second camera 112, as well as a 3D model or mesh of the environment 300, the pass-through image presented to the user 100 may describe objects located in their actual location (i.e., at the appropriate depth values) within the environment 300 from the perspective of the user 100, thereby enabling the user 100 to pick up an object (e.g., a controller, a glass of water, etc.). However, in some cases, image data from the first camera 110 and / or the second camera 112 may be captured continuously, and upon detecting a trigger event, the HMD 104 may display content associated with the environment 300 to the user 100.

[0059] 3 , a triggering event may include determining that user 100 is approaching or approaching a boundary of environment 300. For example, user 100 may approach a corner 304 between two walls 306 and 308 of environment 300. By knowing the location of user 100 via tracking system 122, HMD 104 can display content on display 108 to indicate that user 100 is approaching corner 304. For example, as shown, an indication of corner 304 and walls 306, 308 may be displayed as dashed or dotted lines on display 108. In doing so, user 100 is presented with an indication that user 100 is about to run into a wall. In some cases, as shown in FIG. 3, instructions or content regarding corner 304 and walls 306, 308 may be overlaid (in combination with virtual content displayed on display 108 of HMD 104) or presented on (or in combination with) virtual content 302.

[0060] As described above, image data captured by first camera 110 and / or second camera 112 can be modified to take into account the viewpoint of user 100. Additionally, or alternatively, the depth map, 3D mesh, or model of the environment takes into account or can take into account the location of first camera 110 and / or second camera 112 relative to the eyes or viewpoint of user 100. Thus, if first camera 110 and / or second camera 112 have a viewpoint that may differ from that of user 100, the image data can be modified to adjust to the user's viewpoint before presentation on HMD 104.

[0061] 3 illustrates a particular implementation of displaying the corner 304 and / or the walls 306, 308 on the display 108, the content may be presented differently, such as unmodified, modified for display, embedded, merged, stacked, split, re-rendered, or otherwise manipulated, so as to be appropriately provided to the user 100 without interrupting the immersive virtual experience. For example, in some cases, content may be displayed in a particular region of the display 108 (e.g., the top right-hand corner). Additionally or alternatively, the display 108 may present only content associated with the corner 304 and the walls 306, 308, so as to have 100 percent pass-through. In some cases, the HMD 104 may fade in content associated with the environment 300 on the display 108.

[0062] FIG. 4 illustrates a user 100 wearing an HMD 104 within an environment 400. In some cases, the HMD 104 may display a pass-through image based at least in part on one or more trigger events. In some cases, the trigger event may include determining that a visitor 402 has entered the environment 400 and / or a defined area within the environment 400. As shown in FIG. 4 , the display 108 depicts the visitor 402 approaching the user 100 wearing the HMD 104. Here, the visitor 402 appears to be approaching the user 100, and therefore the display 108 may display images captured by the first camera 110 and / or the second camera 112. However, the images may first be modified to account for differences in perspective between the first camera 110 and / or the second camera 112 and the perspective of the user 100. In some cases, the HMD 104 may display the visitor 402 based at least in part on the visitor 402 being within the user's 100's line of sight and / or in front of the user 100, as determined using the marker(s) 134, while wearing the HMD 104 or while within a threshold distance of the user 100. In some cases, the visitor 402 may be detected via a motion sensor and analyzed from image data from the first camera 110 and / or the second camera 112, such as via a tracking system 122.

[0063] 5 and 6 illustrate processes according to embodiments of the present application. The processes described herein are described as a collection of blocks in logical flow diagrams representing sequences of operations, some or all of which may be implemented in hardware, software, or a combination thereof. In the context of software, the blocks may represent computer-executable instructions stored on one or more computer-readable media, which, when executed by one or more processors, program the processors to perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as limiting unless otherwise specified. Any number of the described blocks may be combined in any order and / or in parallel to implement a process, or alternative processes, and not all of the blocks need be executed. For discussion purposes, the processes are described with reference to the environments, architectures, and systems described in the examples herein, such as those described with reference to FIGS. 1-4, but the processes may be implemented in a wide variety of other environments, architectures, and systems.

[0064] 5 illustrates an example process 500 for generating a 3D model, depth map, or mesh of environment 102. In some cases, process 500 may be performed by a remote computing resource 142. However, HMD 104, tracking system 122, and / or computer 150 may also perform some or all of process 500.

[0065] At 502, the process 500 may transmit a request to define an area within the environment 102. For example, the remote computing resource 142 may transmit a request to the HMD 104 requesting that the user 100 define an area of ​​the environment 102. In some cases, the area may represent an area in which the user 100 intends to move about while immersed in the VR environment. In response to receiving the request, the processor(s) 114 of the HMD 104 may present the request or information associated with the request on the display 108. For example, the request may inform the user 100 to wear the HMD 104, walk around the environment 102, and define an area in which the user intends to move about.

[0066] At 504, process 500 may transmit a request to track the HMD 104 within the environment 102 and as the user 100 wears the HMD 104 and defines a zone. For example, the remote computing resource 142 may transmit a request to the tracking system 122 to track the user 100 throughout the environment 102 while the user 100 defines a zone. In some cases, the tracking system 122 may track the user 100 via projector(s) 130 that project images into the environment 102 and sensor(s) 132 that capture images of light reflected through marker(s) 134 on the HMD 104. In some cases, the remote computing resource 142 may transmit the request of 502 and the request of 504 simultaneously or substantially simultaneously.

[0067] At 506, process 500 may receive first image data from tracking system 122. For example, remote computing resource 142 may receive the first image data from tracking system 122, where the first image data represents the location of HMD 104 and / or user 100 as user 100 walks around environment 102 to define a zone. For example, the first image data may represent the location and / or pose of marker(s) 134 within environment 102. The first image data received by remote computing resource 142 may include a timestamp of when the image was captured by tracking system 122 (or sensor(s) 132).

[0068] At 508, process 500 may determine characteristics of the area. For example, based at least in part on receiving the first image data, remote computing resource 142 may determine a boundary or perimeter of the area. Additionally or alternatively, remote computing resource 142 may determine a center or origin of the area based at least in part on the first image data. In some cases, the origin may be defined in 3D space having values ​​(0, 0, 0) corresponding to X, Y, and Z coordinates, respectively, in a Cartesian coordinate system. Thus, upon receiving subsequent image data from tracking system 122 and / or remote computing resource 142, the relative location of user 100 (and / or HMD 104) with respect to the origin of the area can be determined.

[0069] At 510, process 500 may receive second image data from HMD 104. For example, remote computing resource 142 may receive the second image data from HMD 104, where the second image data represents images captured by first camera 110 and / or second camera 112 while user 100 defined an area of ​​environment 102. In some cases, the second image data received by remote computing resource 142 may include a timestamp of when the images were captured by first camera 110 and / or second camera 112. That is, while first camera 110 and / or second camera 112 capture images, another communicatively coupled computing device may determine the pose of user 100 within environment 102. For example, as described above, while the first camera 110 and / or the second camera 112 capture images, the processor(s) 136 of the tracking system 122 may cause the projector(s) 130 to project images onto the environment 102. The marker(s) 134 of the HMD 104 may reflect light associated with the image, and the sensor(s) 132 may capture reflected images of the marker(s) 134. Such images may be used to determine the depth, location, orientation, and / or pose of the user 100 (and / or HMD 104) within the environment 102 and may be associated with the image data captured by the first camera 110 and / or the second camera 112.

[0070] At 512, the process 500 may generate a 3D model (or mesh) of the environment 102. For example, the remote computing resource 142 may generate the 3D model of the environment 102 based at least in part on the first image data and / or the second image data. For example, using the second image data, the remote computing resource 142 may compare images captured by the first camera 110 and / or the second camera 112, respectively, to determine disparity. The remote computing resource 142 may then determine depth information for the environment 102 for use in generating the 3D model of the environment 102. Additionally, the remote computing resource 142 may utilize the first image data received from the tracking system 122 to associate a depth map or depth values ​​of the environment 102 with areas and / or specific locations within the environment 102. In some cases, the remote computing resource 142 may compare or correlate the first and second image data (or depth maps generated from the image data) using timestamps of when the first and / or second image data were captured. In some cases, generating the 3D model of the environment 102 may include depicting objects in the environment. For example, the remote computing resource 142 may augment the VR environment with objects (or volumes) in the environment 102, such as a chair. In some cases, the 3D model may be transmitted to the HMD 104 and / or the tracking system 122 and / or stored in memory of the HMD 104, the tracking system 122, and / or the remote computing resource 142.

[0071] 6 illustrates an example process 600 for pass-through of an image on an HMD, such as HMD 104. In some cases, process 600 may continue from 512 of process 500 after a 3D model of environment 102 has been generated. In other words, in some cases, after the 3D model has been generated, HMD 104 may display a pass-through image to enable user 100 to switch between the VR environment and the real-world environment (i.e., environment 102). In some cases, process 600 may be performed by remote computing resource 142. However, HMD 104, tracking system 122, and / or computer 150 may also perform some or all of process 600.

[0072] At 602, process 600 may receive first image data from HMD 104. For example, remote computing resource 142 may receive the first image data from HMD 104, where the first image data represents images captured by first camera 110 and / or second camera 112. In some cases, remote computing resource 142 may receive the first image data based, at least in part, on HMD 104 detecting a trigger event, such as user 100 pressing a button, user 100 issuing a verbal command, motion detected within environment 102 (e.g., a visitor approaching user 100), and / or user 100 approaching or reaching within a threshold distance of a boundary of environment 102. In other examples, remote computing resource 142 may receive the first image data continuously and may be configured to pass through these images based, at least in part, on a trigger event. For example, as the user 100 may be immersed in the VR environment, the user 100 may press a button on the HMD 104 to display content outside of the HMD 104. In this sense, the HMD 104 may include a pass-through mode that displays images captured by the first camera 110 and / or the second camera 112. Thus, based at least in part on detecting a triggering event, the processor(s) 114 of the HMD 104 may cause the first camera 110 and / or the second camera 112 to capture images of the environment 102.

[0073] At 604, process 600 may generate a depth map and / or a 3D mesh based at least in part on the first image data. For example, remote computing resource 142 may generate the depth map and / or the 3D mesh based at least in part on the first image data being received from HMD 104 (i.e., using stereo camera image processing). If the first image data represents a portion of environment 102, the depth map and / or the 3D mesh may also correspond to a depth map and / or a 3D mesh of the portion of environment 102. Additionally or alternatively, HMD 104 may generate the depth map and / or the 3D mesh. For example, upon receiving image data from first camera 110 and / or second camera 112, processor(s) 114 may utilize stereoscopic camera image processing to generate the depth map. In some cases, using the depth map, processor(s) 114 may generate a 3D mesh of environment 102. Additionally, in some cases, the processor(s) 114 may store depth maps and / or 3D meshes, such as depth map 118 and / or 3D mesh 120, in memory 116.

[0074] In some cases, the HMD 104, the remote computing resource 142, and / or the computer 150 may modify the first image data received from the first camera 110 and / or the second camera 112 before generating the depth map and / or the 3D mesh. For example, if the first image data cannot represent the viewpoint of the user 100, the first image data may be modified (e.g., repurposed, transformed, distorted, etc.) to take into account differences between the viewpoints of the first camera 110 and the second camera 112 and the viewpoint of the user 100. In some cases, the viewpoint of the user 100 may be determined using the tracking system 122 and the positions of the marker(s) 134 in the environment 102. Additionally or alternatively, after generating the depth map and / or the 3D mesh, the depth map and / or the 3D mesh may be modified according to or based at least in part on the viewpoint of the user 100.

[0075] At 606, process 600 may receive second image data forming tracking system 122 that represents the pose of HMD 104. For example, remote computing resource 142 may receive second image data corresponding to HMD 104 in environment 102 from tracking system 122. As discussed above, the second image data may be captured by sensor(s) 132 of tracking system 122 that detects light reflected by marker(s) 134 of HMD 104 in response to images projected by projector(s) 130.

[0076] At 608, process 600 may determine a pose of the HMD 104. For example, based at least in part on receiving the second image data, the remote computing resource 142 may determine a pose of the HMD 104 within the environment 102. In some cases, the pose may represent a location of the HMD 104 within the environment 102 and / or an orientation of the HMD 104 within the environment 102. That is, the remote computing resource 142 may analyze the first image data and / or the second image data to determine a location of the user 100 within the environment 102 relative to a center of the environment 102. Such analysis may determine a relative location, line of sight, and / or viewpoint of the user 100 relative to the center of the environment 102 or a particular area within the environment 102.

[0077] At 610, process 600 may convert the depth map and / or 3D mesh into points associated with a 3D model of the environment. For example, the remote computing resource 142 can utilize the pose (e.g., location and orientation) to determine the location of the user 100 within the environment 102 and / or the viewpoint of the user 100 within the environment 102. The remote computing resource 142 can convert the points of the depth map and / or 3D mesh into points associated with the 3D model of the environment 102, thereby taking the viewpoint of the user 100 into account. In other words, the processor(s) 144 of the remote computing resource 142 can locate, find, or determine depth values ​​of points within the 3D model of the environment 102 by converting the points of the depth map and / or 3D mesh into points associated with the 3D model of the environment 102. Using the pose, the remote computing resource 142 can transfer the depth map and / or 3D mesh generated at 604 into the 3D model of the environment 102. Such repurposing can help to accurately depict the environment 102 (e.g., at the proper depth value) relative to the user 100. That is, the first image data, together with the second image data and / or a 3D model of the environment 102, can provide the absolute position and / or line of sight of the HMD 104, which can help to depict the viewpoint of the user 100 at the proper depth value or depth perception.

[0078] At 612, the process 600 can project the first image data onto a portion of a 3D model to generate third image data. For example, with the pose of the user 100 known, the remote computing resource 142 can project or overlay the first image data onto a portion of a 3D model of the environment 102 to generate the third image data. In some cases, generating the third image data can include cross-blending or inputting depth and / or color values ​​for particular pixels of the third image data. For example, if the first image data cannot represent the viewpoint of the user 100, when generating the third image data depicting the viewpoint of the user 100, the third image data can have undefined depth and / or color values ​​for particular pixels. Here, pixels that do not include color values ​​can be assigned color values ​​from neighboring pixels or an average value of the neighboring pixels. Additionally or alternatively, the third image data can be generated using a previous depth map or 3D mesh of the environment 100, particle system, etc.

[0079] At 614, process 600 may transmit third image data to HMD 104. For example, after projecting the first image data onto the portion of the 3D model, remote computing resource 142 may transmit third image data to HMD 104, the third image data representing the first image data projected onto the portion of the 3D model.

[0080] From 614, process 600 can loop back to 602 to receive subsequent image data. As a result, in response to a continuing trigger event (e.g., a button press, a voice command, motion detection, etc.), the HMD 104 can transition to pass-through mode, providing a convenient way for the user 100 to leave the environment 102 without having to remove the HMD 104. For example, the remote computing resource 142 may receive an indication from the tracking system 122 that the user 100 is approaching an area boundary and / or is about to collide with a wall of the environment 102. Upon receiving this indication, the remote computing resource 142 may receive image data from the HMD 104 representing the user's perspective (i.e., images captured by the first camera 110 and / or the second camera 112). Upon determining the user's pose, the remote computing resource 142 can project the image data onto a 3D model of the environment 102 and display the image data on the HMD 104. In this sense, images may be automatically "passed through" to the user 100, allowing the user to view the real-world environment without having to break immersion.

[0081] While the above invention has been described with reference to specific embodiments, it should be understood that the scope of the invention is not limited to these specific embodiments. Since other modifications and variations adapted to suit particular operating requirements and environments will be apparent to those skilled in the art, the present invention is not to be regarded as limited to the embodiments selected for purposes of disclosure, but extends to all variations and variations that do not depart from the true spirit and scope of the invention.

[0082] Although the present application describes embodiments having particular structural features and / or methodological acts, it should be understood that the claims are not necessarily limited to the particular features or acts described. Rather, the particular features and acts are merely illustrative of some embodiments that may be within the scope of the present application's claims.

Claims

1. display, a first camera, and a head mounted display including a second camera; one or more processors; one or more non-transitory computer-readable media storing computer-executable instructions; the computer-executable instructions, when executed by the one or more processors, cause the one or more processors to perform actions, the actions including: capturing first image data representing a first portion of an environment via the first camera; capturing second image data representative of a second portion of the environment via the second camera; determining a first offset of the first camera relative to a first eye of a user and a second offset of the second camera relative to a second eye of the user; generating a depth map based at least in part on the first image data, the second image data, the first offset, and the second offset; transmitting the depth map, the first image data, and the second image data to one or more computing devices; receiving third image data from the one or more computing devices based at least in part on the depth map, the first image data, and the second image data; and displaying the third image data via the display.

2. the head-mounted display further comprises a button; the first image data is captured based at least in part on detecting a depression of the button; or The system of claim 1 , wherein the second image data is captured based at least in part on detecting the depression of the button.

3. the one or more non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform actions, the actions further including receiving an indication from the one or more computing devices that the user is approaching a boundary of the environment; capturing the first image data is based at least in part on receiving the instruction; or The system of claim 1 , wherein capturing the second image data is at least one of based at least in part on receiving the instruction.

4. the users include a first user, the one or more non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform actions, the actions further including receiving an indication from the one or more computing devices that a second user is present in the environment; capturing the first image data is based at least in part on receiving the instruction; or The system of claim 1 , wherein capturing the second image data is at least one of based at least in part on receiving the instruction.

5. the third image data is displayed in combination with virtual content presented on the display; or The system of claim 1 , wherein the third image data is at least one of displayed on a predetermined portion of the display.

6. 1. A method comprising: receiving first image data representing at least a portion of an environment from a head mounted display; receiving second image data from a tracking system indicating a viewpoint of a user within the environment; generating a depth map corresponding to at least a portion of the environment based at least in part on at least one of the first image data or the second image data; Associating the depth map with a portion of a three-dimensional (3D) model of the environment; transposing the first image data to the portion of the 3D model of the environment; generating third image data representing the first image data transposed onto the portion of the 3D model of the environment; transmitting the third image data to the head mounted display.

7. transmitting a first request to the head mounted display to capture an image of the environment; 7. The method of claim 6, further comprising transmitting a second request to the tracking system to track a location of the head mounted display while the head mounted display is capturing the image of the environment.

8. receiving fourth image data from the head mounted display representing the image of the environment; receiving fifth image data from the tracking system while capturing the image of the environment, the fifth image data representing the location of the head mounted display; The method of claim 7 , further comprising: generating a 3D model of the environment based at least in part on the fourth image data and the fifth image data.

9. The method of claim 6 , wherein associating the depth map with a portion of the 3D model of the environment comprises aligning individual points of the depth map with corresponding individual points of the portion of the 3D model.

10. The method of claim 6 , wherein transmitting the third image data is based at least in part on receiving an instruction from the head-mounted display.

11. The instructions are: a button is pressed on the head-mounted display; detecting motion in the environment; detecting movement of said user within a predetermined distance; or The method of claim 10 , wherein the method corresponds to at least one of determining that the user is approaching a boundary of the environment.

12. 7. The method of claim 6, further comprising transmitting fourth image data to the head-mounted display, the fourth image data representing virtual content associated with a virtual environment, and the head-mounted display configured to display at least a portion of the third image data in combination with at least a portion of the fourth image data.

13. determining a center of the environment; determining a location of the user within the environment based at least in part on the second image data; The method of claim 6 , wherein associating the depth map with the portion of the 3D model is based at least in part on determining the location of the user relative to the center of the environment.

14. The method of claim 6 , wherein the third image data represents the viewpoint of the user within the environment.

15. 1. A system comprising: A head-mounted display and A first camera; A second camera; one or more processors; and one or more non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform actions, including: receiving first image data representing a first portion of an environment from the first camera; receiving second image data from the second camera representing a second portion of the environment; generating a first depth map of a portion of the environment based at least in part on the first image data and the second image data; receiving data from a tracking system corresponding to a user's viewpoint within the environment; generating a second depth map based at least in part on the first depth map and the data; generating third image data based at least in part on projecting the first image data and the second image data onto the second depth map; transmitting the third image data to the head-mounted display.

16. 16. The system of claim 15, wherein the one or more non-transitory computer-readable media store computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform actions, the actions further including receiving an indication of a trigger event, and transmitting the third image data is based at least in part on receiving the indication.

17. The instructions are: determining that another user is in front of the user wearing the head mounted display; or The system of claim 16 , responsive to at least one of determining that the user is approaching a boundary of the environment.

18. 16. The system of claim 15, wherein the one or more non-transitory computer-readable media store computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform an action, the action further comprising transmitting fourth image data representing virtual content within a virtual reality environment.

19. The system of claim 15 , wherein the data corresponds to a location of the user within the environment relative to a central location of the environment.

20. The one or more non-transitory computer-readable media store computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform actions, including: generating a 3D model of the environment; determining a portion of the 3D model that corresponds to the viewpoint of the user; The system of claim 15 , wherein generating the second depth map is based at least in part on the portion of the 3D model.