Non-uniform Stereo Rendering

By aligning and cropping virtual environment images to match the real environment's center and adjusting aspect ratios and resolutions, MR recording systems achieve improved integration and immersion in shared experiences.

JP7762780B2Active Publication Date: 2025-10-30MAGIC LEAP INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024145629
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-25
Filing Date
2024-08-27
Publication Date
2025-10-30
Estimated Expiration
2040-10-23

AI Technical Summary

Technical Problem

Existing mixed reality (MR) recording systems struggle to accurately integrate real and virtual environments due to differences in line of sight and resolution, resulting in poor integration and immersion when sharing MR experiences.

Method used

A method and system for generating and presenting a combined image by estimating the pose of a wearable head device, creating a virtual environment image with a larger field of view than the real environment, and aligning and cropping it to match the real environment's center, ensuring synchronized aspect ratios and resolutions.

Benefits of technology

This approach enables an immersive, integrated recording and display of MR experiences that maintain the user's first-person perspective, enhancing the sense of presence and reducing visual discrepancies between real and virtual content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007762780000001
    Figure 0007762780000001
  • Figure 0007762780000002
    Figure 0007762780000002
  • Figure 0007762780000003
    Figure 0007762780000003
Patent Text Reader

Abstract

To provide non-uniform stereo rendering.SOLUTION: Embodiments of the present disclosure describe systems and methods for recording augmented reality and mixed reality experiences. In an exemplary method, an image of a real world environment is received via a camera of a wearable head device. A posture of the wearable head device is estimated, and a first image of the virtual environment is generated based on the posture. A second image of the virtual environment is generated based on the posture, and the second image of the virtual environment has a field of view larger than that of the first image of the virtual environment. A combined image is generated based on the second image of the virtual environment and the image of the real environment.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Application No. 62 / 926,306, filed October 25, 2019, the entire disclosure of which is incorporated herein by reference for all purposes.

[0002] The present disclosure relates generally to systems and methods for rendering and displaying visual information, and more particularly to systems and methods for rendering and displaying visual information within a mixed reality environment. [Background technology]

[0003] Virtual environments are ubiquitous in computing environments, finding use in video games (where a virtual environment may represent a game world), maps (where a virtual environment may represent a terrain to be navigated), simulations (where a virtual environment may simulate a real environment), digital storytelling (where virtual characters may interact with one another within a virtual environment), and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, a user's experience with a virtual environment may be limited by the technology for presenting the virtual environment. For example, traditional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may be unable to realize a virtual environment in a way that creates a compelling, realistic, and immersive experience.

[0004] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present to a user of an XR system sensory information corresponding to a virtual environment represented by data in a computer system. This disclosure considers uniqueness among VR, AR, and MR systems (although some systems may be categorized as VR in one aspect (e.g., visual aspect) and simultaneously categorized as AR or MR in another aspect (e.g., audio aspect)). As used herein, a VR system presents a virtual environment that replaces the user's real environment in at least one aspect. For example, a VR system may present a user with a view of the virtual environment while simultaneously obscuring that view of the real environment, such as with an optically blocking head-mounted display. Similarly, a VR system may present a user with audio corresponding to the virtual environment while simultaneously blocking (attenuating) the audio from the real environment.

[0005] VR systems may suffer from various drawbacks resulting from replacing a user's real environment with a virtual environment. One drawback is motion sickness, which can occur when a user's field of view within the virtual environment no longer corresponds to the state of their inner ear, which detects their balance and orientation in the real (but not the virtual) environment. Similarly, a user may experience disorientation within a VR environment if their body and limbs (the view upon which the user relies to feel "grounded" in the real environment) are not directly visible. Another drawback is the computational burden (e.g., memory, processing power) imposed on a VR system that must present a fully 3D virtual environment, especially in real-time applications that seek to immerse a user in the virtual environment. Similarly, such an environment may need to reach a very high level of realism to be considered immersive, as users tend to be sensitive to even slight imperfections in the virtual environment, any of which can destroy the user's sense of immersion in the virtual environment. Furthermore, another disadvantage of VR systems is that such applications of the systems cannot take advantage of the wide range of sensory data in the real environment, such as the various sights and sounds experienced in the real world. A related disadvantage is that VR systems may struggle to create shared environments in which multiple users can interact, because users who share physical space in the real environment may not be able to see or interact with each other directly in the virtual environment.

[0006] As used herein, an AR system presents a virtual environment that overlaps or overlays the real environment in at least one aspect. For example, an AR system may present a user with a view of the virtual environment overlaid on the user's view of the real environment, such as using a see-through head-mounted display that presents a displayed image while allowing light to pass through the display into the user's eyes. Similarly, an AR system may present a user with audio corresponding to the virtual environment while simultaneously mixing in audio from the real environment. Similarly, as used herein, an MR system, like an AR system, may present a virtual environment that overlaps or overlays the real environment in at least one aspect, and may additionally allow the virtual environment in the MR system to interact with the real environment in at least one aspect. For example, a virtual character in the virtual environment may flip a light switch in the real environment, causing a corresponding light bulb in the real environment to turn on or off. As another example, the virtual character may react to audio signals in the real environment (such as with facial expressions). By maintaining the presentation of the real environment, AR and MR systems may avoid some of the aforementioned disadvantages of VR systems. For example, motion sickness in a user is reduced because visual cues from the real environment (including the user's own body) can remain visible and such systems do not need to present the user with a fully realized 3D environment to be immersive. Furthermore, AR and MR systems can create new applications that utilize real-world sensory input (e.g., views and sounds of scenery, objects, and other users) to augment that input.

[0007] It may be desirable to capture and record MR experiences so that they can be shared with other users. MR systems (especially wearable head devices) are better positioned than simple video recording systems and can record environments (both real and virtual) in the same way that a user experiences the environment. Providing a firsthand view into a user's experience within real and virtual environments can make the recording feel more immersive, as if the viewer were with the original user at the time of recording. Recording an MR experience may include combining recordings of a real environment with recordings of a virtual environment. Recorded MR content may be useful for both casual content sharing and expert content creation. For example, it may be desirable to share a casual MR experience (e.g., a concert) with other users, who may view the recorded MR content on a two-dimensional screen or on an MR system. It may also be desirable for expert content creators to record MR content (e.g., advertisements) for display on other MR systems or two-dimensional screens. However, recording an MR experience may not be as trivial as recording a simple two-dimensional video, at least because the camera used to record the real environment may be positioned offset from the user's eyes so as to obstruct the user's view. However, the virtual content rendered for the user may be rendered from the line of sight of the user's eyes. This offset may create a difference in line of sight between the object captured by the camera and the virtual content rendered for the user when combining the two recording streams. Furthermore, the camera may record at a different resolution and / or aspect ratio than the displayed virtual content. Combining recordings of the real environment and the virtual environment may therefore result in poor integration between the real and virtual content (e.g., the real and virtual content may be presented from different line of sight). Therefore, it may be desirable to create an accurate MR recording system that simulates a first-person experience and displays the recorded MR content in an integrated and accurate manner. Summary of the Invention [Means for solving the problem]

[0008]

[0003] Examples of the present disclosure describe systems and methods for recording augmented reality and mixed reality experiences. In an exemplary method, an image of a real environment is received via a camera of a wearable head device. A pose of the wearable head device is estimated, and a first image of a virtual environment is generated based on the pose. A second image of the virtual environment is generated based on the pose, the second image of the virtual environment having a field of view larger than the field of view of the first image of the virtual environment. A combined image is generated based on the second image of the virtual environment and the image of the real environment. The present invention provides, for example, the following items. (Item 1) 1. A method comprising: receiving an image of a real environment via a camera of a wearable head device; estimating a pose of the wearable head device; generating a first image of a virtual environment based on the pose; generating a second image of the virtual environment based on the pose, the second image of the virtual environment having a field of view that is larger than a field of view of the first image of the virtual environment; generating a combined image based on the second image of the virtual environment and the image of the real environment; A method comprising: (Item 2) generating a cropped image based on the second image of the virtual environment; presenting a stereoscopic image based on the first image of the virtual environment and the cropped image via a display of the wearable head device; and Item 1, the method of claim 1 further comprising: (Item 3) Item 3. The method of item 2, wherein the cropped image comprises a center of the cropped image, and the second image of the virtual environment comprises a center of the second image, and the center of the cropped image is the same as the center of the second image. (Item 4) Item 3. The method of item 2, wherein the cropped image comprises a cropped field of view, and a field of view of the second image of the virtual environment is larger than the cropped field of view. (Item 5) Item 10. The method of item 1, wherein generating the combined image comprises planar projection. (Item 6) Item 10. The method of item 1, wherein the second image of the virtual environment has an aspect ratio and the image of the real environment has the same aspect ratio. (Item 7) Item 10. The method of item 1, wherein the second image of the virtual environment has an aspect ratio and the first image of the virtual environment has the same aspect ratio. (Item 8) Item 10. The method of claim 1, wherein a first image of the virtual environment has a first frame buffer size and a second image of the virtual environment has a second frame buffer size, the second frame buffer size being larger than the first frame buffer size. (Item 9) Item 10. The method of item 1, wherein the first image of the virtual environment has a frame buffer size and the cropped image has the same frame buffer size. (Item 10) Item 10. The method of item 1, wherein the first image of the virtual environment has a first resolution and the second image of the virtual environment has a second resolution, the second resolution being greater than the first resolution. (Item 11) Item 10. The method of item 1, wherein the camera is configured to be placed to the left of the user's left eye. (Item 12) Item 10. The method of item 1, wherein the camera is configured to be placed to the right of the user's right eye. (Item 13) Item 10. The method of item 1, wherein the camera is configured to be placed between a user's left eye and a user's right eye. (Item 14) Item 10. The method of item 1, further comprising presenting the combined image via a display on a second device different from the wearable head device. (Item 15) 1. A system comprising: A camera on a wearable head device, one or more processors, receiving an image of a real environment via a camera of the wearable head device; estimating a pose of the wearable head device; generating a first image of a virtual environment based on the pose; generating a second image of the virtual environment based on the pose, the second image of the virtual environment having a field of view that is larger than a field of view of the first image of the virtual environment; generating a combined image based on the second image of the virtual environment and the image of the real environment; one or more processors configured to perform a method including: A system comprising: (Item 16) generating a cropped image based on the second image of the virtual environment; presenting a stereoscopic image based on the first image of the virtual environment and the cropped image via a display of the wearable head device; and Item 16. The system of item 15, further comprising: (Item 17) Item 17. The system of item 16, wherein the cropped image comprises a center of the cropped image, and the second image of the virtual environment comprises a center of the second image, and the center of the cropped image is the same as the center of the second image. (Item 18) A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receiving an image of a real environment via a camera of a wearable head device; estimating a pose of the wearable head device; generating a first image of a virtual environment based on the pose; generating a second image of the virtual environment based on the pose, the second image of the virtual environment having a field of view that is larger than a field of view of the first image of the virtual environment; generating a combined image based on the second image of the virtual environment and the image of the real environment; A non-transitory computer-readable medium for performing a method comprising: (Item 19) generating a cropped image based on the second image of the virtual environment; presenting a stereoscopic image based on the first image of the virtual environment and the cropped image via a display of the wearable head device; and Item 19. The non-transitory computer-readable medium of item 18, further comprising: (Item 20) 20. The non-transitory computer-readable medium of claim 19, wherein the cropped image comprises a center of the cropped image, and the second image of the virtual environment comprises a center of the second image, and the center of the cropped image is the same as the center of the second image. [Brief explanation of the drawings]

[0009] [Figure 1A] 1A-1C illustrate an exemplary mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 1B] 1A-1C illustrate an exemplary mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 1C] 1A-1C illustrate an exemplary mixed reality environment in accordance with one or more embodiments of the present disclosure.

[0010] [Figure 2A] 2A-2D illustrate components of an example mixed reality system that may be used to generate and interact with a mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 2B] 2A-2D illustrate components of an example mixed reality system that may be used to generate and interact with a mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 2C] 2A-2D illustrate components of an example mixed reality system that may be used to generate and interact with a mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 2D] 2A-2D illustrate components of an example mixed reality system that may be used to generate and interact with a mixed reality environment in accordance with one or more embodiments of the present disclosure.

[0011] [Figure 3A] FIG. 3A illustrates an example mixed reality handheld controller that may be used to provide input to a mixed reality environment, according to one or more embodiments of the present disclosure.

[0012] [Figure 3B] FIG. 3B illustrates an example auxiliary unit that may be used in conjunction with an example mixed reality system in accordance with one or more embodiments of the present disclosure.

[0013] [Figure 4] FIG. 4 illustrates an example functional block diagram for an example mixed reality system in accordance with one or more embodiments of the present disclosure.

[0014] [Figure 5] 5A-5B illustrate an example of mixed reality recording according to one or more embodiments of the present disclosure.

[0015] [Figure 6] FIG. 6 illustrates an example of mixed reality recording, according to one or more embodiments of the present disclosure.

[0016] [Figure 7] FIG. 7 illustrates an example of mixed reality recording, according to one or more embodiments of the present disclosure.

[0017] [Figure 8] FIG. 8 illustrates an example of mixed reality recording, according to one or more embodiments of the present disclosure.

[0018] [Figure 9] FIG. 9 illustrates an example of mixed reality recording, according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0019] Detailed Description In the following description of the embodiments, reference is made to the accompanying drawings which form a part hereof, and in which is shown, by way of illustration, specific embodiments which may be practiced. It is to be understood that other embodiments may be used and structural changes may be made without departing from the scope of the disclosed embodiments.

[0020] Mixed Reality Environment

[0021] Like all people, users of mixed reality systems exist in a real environment, i.e., three-dimensional portions of the "real world" and all of its content are perceptible to the user. For example, users perceive the real environment using normal human senses, i.e., sight, hearing, touch, taste, and smell, and interact with the real environment by moving their body within the real environment. Locations within the real environment can be described as coordinates within a coordinate space. For example, coordinates can include latitude, longitude, and altitude relative to sea level, distance in three orthogonal dimensions from a reference point, or other suitable values. Similarly, a vector can describe a quantity, having a direction and magnitude within the coordinate space.

[0022] A computing device may maintain a representation of a virtual environment, for example, in a memory associated with the device. As used herein, a virtual environment is a computed representation of a three-dimensional space. The virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other properties associated with that space. In some examples, circuitry (e.g., a processor) of a computing device may maintain and update the state of the virtual environment. That is, the processor may determine the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or input provided by a user at a first time t0. For example, if an object in the virtual environment is located at a first coordinate and has certain programmed physical parameters (e.g., mass, coefficient of friction) at time t0, and input received from the user indicates that a force should be applied to the object in a certain directional vector, the processor may apply the laws of kinematics and use basic mechanics to determine the location of the object at time t1. The processor may determine the state of the virtual environment at time t1 using any suitable information known about the virtual environment and / or any suitable input. In maintaining and updating the state of the virtual environment, the processor may execute any suitable software, including software related to creating and deleting virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling input and output, software for implementing network operations, software for applying asset data (e.g., animation data for moving a virtual object over time), or many other possibilities.

[0023] An output device, such as a display or speakers, can present any or all aspects of the virtual environment to the user. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, lights, etc.) that can be presented to the user. A processor can determine a view of the virtual environment (e.g., corresponding to a "camera," with its origin coordinates, viewing axis, and frustum) and render on the display a viewable scene of the virtual environment corresponding to that view. Any suitable rendering technique may be used for this purpose. In some examples, the viewable scene may include only some virtual objects in the virtual environment and exclude certain other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, a virtual object in the virtual environment may generate a sound originating from the object's location coordinates (e.g., a virtual character may speak or produce a sound effect), or the virtual environment may be associated with a musical cue or ambient sound that may or may not be associated with a particular location. The processor can determine audio signals corresponding to the "listener" coordinates, e.g., audio signals corresponding to the synthesis of sounds in the virtual environment and mixed and processed to simulate the audio signals that would be heard by a listener at the listener coordinates, and present the audio signals to the user via one or more speakers.

[0024] Because the virtual environment exists only as a computational construct, the user cannot directly perceive the virtual environment using their normal senses. Instead, the user can only indirectly perceive the virtual environment, as presented to the user, for example, by a display, speakers, tactile output device, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data via input devices or sensors to a processor, which can use the device or sensor data to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use that data to cause the object to respond appropriately within the virtual environment.

[0025] A mixed reality system can present a user with a mixed reality environment (“MRE”) that combines aspects of a real environment and a virtual environment, for example, using a see-through display and / or one or more speakers (which may, for example, be incorporated into a wearable head device). In some embodiments, the one or more speakers may be external to the head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of a real environment and a corresponding virtual environment. In some examples, the corresponding real and virtual environments share a single coordinate space. In some examples, the real coordinate space and the corresponding virtual coordinate space are related to each other by a transformation matrix (or other suitable representation). Thus, a single coordinate (in some examples, together with the transformation matrix) may define a first location in the real environment and a second corresponding location in the virtual environment, and vice versa.

[0026] In an MRE, a virtual object (e.g., in a virtual environment associated with the MRE) may correspond to a real object (e.g., in a real environment associated with the MRE). For example, if the real environment of the MRE includes a real lamppost (real object) at certain location coordinates, the virtual environment of the MRE may include a virtual lamppost (virtual object) at the corresponding location coordinates. As used herein, a real object combined with its corresponding virtual object constitutes a “mixed reality object.” It is not necessary for a virtual object to perfectly match or match the corresponding real object. In some embodiments, a virtual object can be a simplified version of the corresponding real object. For example, if the real environment includes a real lamppost, the corresponding virtual object may include a cylinder of approximately the same height and radius as the real lamppost (reflecting that a lamppost may be approximately cylindrical in shape). Simplifying virtual objects in this way can enable computational efficiency and simplify calculations to be performed on such virtual objects. Furthermore, in some embodiments of an MRE, not all real objects in the real environment may be associated with a corresponding virtual object. Similarly, in some embodiments of an MRE, not all virtual objects in the virtual environment may be associated with corresponding real objects, i.e., some virtual objects may exist solely within the virtual environment of the MRE without any real-world counterpart.

[0027] In some embodiments, virtual objects may have characteristics that differ, sometimes significantly, from those of their corresponding real objects. For example, a real environment in an MRE may include a green, two-pronged cactus, i.e., a thorny, inanimate object, while the corresponding virtual object in the MRE may have the characteristics of a green, two-armed virtual character with human facial features and a surly attitude. In this embodiment, the virtual object resembles its corresponding real object in some characteristics (color, number of arms) but differs from the real object in other characteristics (facial features, personality). In this manner, virtual objects have the potential to represent real objects in a creative, abstract, exaggerated, or fictional manner, or to impart behavior (e.g., human personality) to otherwise inanimate real objects. In some embodiments, a virtual object may be a purely fictional creation with no real-world counterpart (e.g., a virtual monster in a virtual environment, perhaps in a location that corresponds to a void in the real environment).

[0028] Compared to VR systems, which present a virtual environment to a user while obscuring the real environment, mixed reality systems that present an MRE offer the advantage that the real environment remains perceptible while the virtual environment is presented. Thus, a user of a mixed reality system can experience and interact with the corresponding virtual environment using visual and audio cues associated with the real environment. As an example, a user of a VR system may struggle to perceive or interact with virtual objects displayed in the virtual environment because, as noted above, the user cannot directly perceive or interact with the virtual environment. However, a user of an MR system may find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching the corresponding real objects in their own real environment. This level of interaction may enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by simultaneously presenting a real environment and a virtual environment, a mixed reality system may reduce negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. Mixed reality systems also offer many possibilities for applications that can augment or modify our experience of the real world.

[0029] 1A illustrates an exemplary real environment 100 in which a user 110 uses a mixed reality system 112. The mixed reality system 112 may include a display (e.g., a see-through display) and one or more speakers, as well as one or more sensors (e.g., cameras), for example, as described below. The illustrated real environment 100 includes a rectangular room 104A in which the user 110 is standing and real objects 122A (lamp), 124A (table), 126A (sofa), and 128A (painting). The room 104A further includes a location coordinate 106, which may be considered the origin of the real environment 100. As shown in FIG. 1A, an environment / world coordinate system 108 (comprising an x-axis 108X, a y-axis 108Y, and a z-axis 108Z), with its origin at point 106 (world coordinates), may define a coordinate space for the real environment 100. In some embodiments, origin 106 of environment / world coordinate system 108 may correspond to where mixed reality system 112 is powered on. In some embodiments, origin 106 of environment / world coordinate system 108 may be reset during operation. In some examples, user 110 may be considered a real object in real environment 100. Similarly, body parts (e.g., hands, feet) of user 110 may be considered real objects in real environment 100. In some examples, user / listener / head coordinate system 114 (comprising x-axis 114X, y-axis 114Y, and z-axis 114Z), with its origin at point 115 (e.g., user / listener / head coordinate), may define a coordinate space for user / listener / head on which mixed reality system 112 is located. Origin 115 of user / listener / head coordinate system 114 may be defined relative to one or more components of mixed reality system 112. For example, the origin 115 of the user / listener / head coordinate system 114 may be defined relative to the display of the mixed reality system 112, such as during an initial calibration of the mixed reality system 112. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the user / listener / head coordinate system 114 space and the environment / world coordinate system 108 space.In some embodiments, left ear coordinates 116 and right ear coordinates 117 may be defined relative to the origin 115 of the user / listener / head coordinate system 114. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the left ear coordinates 116 and right ear coordinates 117 and the user / listener / head coordinate system 114 space. The user / listener / head coordinate system 114 can simplify the representation of location relative to the user's head or head-mounted device, for example, relative to the environment / world coordinate system 108. Using simultaneous localization and mapping (SLAM), visual odometry, or other techniques, the transformation between the user coordinate system 114 and the environment coordinate system 108 can be determined and updated in real time.

[0030] 1B illustrates an exemplary virtual environment 130 that corresponds to real environment 100. The illustrated virtual environment 130 includes a virtual rectangular room 104B that corresponds to real rectangular room 104A, a virtual object 122B that corresponds to real object 122A, a virtual object 124B that corresponds to real object 124A, and a virtual object 126B that corresponds to real object 126A. Metadata associated with virtual objects 122B, 124B, and 126B may include information derived from the corresponding real objects 122A, 124A, and 126A. Virtual environment 130 additionally includes a virtual monster 132, which does not correspond to any real object in real environment 100. Real object 128A in real environment 100 does not correspond to any virtual object in virtual environment 130. A persistent coordinate system 133 (with x-axis 133X, y-axis 133Y, and z-axis 133Z), with its origin at point 134 (persistent coordinate), may define a coordinate space for the virtual content. Origin 134 of persistent coordinate system 133 may be defined relative to / with respect to one or more real objects, such as real object 126A. Matrices (which may include translation matrices and quaternion or other rotation matrices) or other suitable representations can characterize the transformation between persistent coordinate system 133 space and environment / world coordinate system 108 space. In some embodiments, virtual objects 122B, 124B, 126B, and 132 may each have its own persistent coordinate point relative to origin 134 of persistent coordinate system 133. In some embodiments, there may be multiple persistent coordinate systems, and virtual objects 122B, 124B, 126B, and 132 may each have its own persistent coordinate point relative to one or more persistent coordinate systems.

[0031] 1A and 1B, environment / world coordinate system 108 defines a shared coordinate space for both real environment 100 and virtual environment 130. In the illustrated embodiment, the coordinate space has its origin at point 106. Furthermore, the coordinate space is defined by the same three orthogonal axes (108X, 108Y, 108Z). Thus, a first location in real environment 100 and a second corresponding location in virtual environment 130 can be described with respect to the same coordinate space. This simplifies identifying and displaying corresponding locations in the real and virtual environments because the same coordinates can be used to identify both locations. However, in some embodiments, the corresponding real and virtual environments need not use a shared coordinate space. For example, in some embodiments (not shown), a matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.

[0032] 1C illustrates an exemplary MRE 150 that simultaneously presents aspects of real environment 100 and virtual environment 130 to user 110 via mixed reality system 112. In the example shown, MRE 150 simultaneously presents to user 110 real objects 122A, 124A, 126A, and 128A from real environment 100 (e.g., through a translucent portion of the display of mixed reality system 112) and virtual objects 122B, 124B, 126B, and 132 from virtual environment 130 (e.g., through an active display portion of mixed reality system 112). As described above, origin 106 serves as the origin for a coordinate space corresponding to MRE 150, and coordinate system 108 defines the x-, y-, and z-axes for the coordinate space.

[0033] In the illustrated example, the mixed reality objects include corresponding pairs of real and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) that occupy corresponding locations in coordinate space 108. In some examples, both real and virtual objects may be visible to user 110 simultaneously. This may be desirable in instances where, for example, a virtual object presents information designed to augment the view of the corresponding real object (such as in a museum application where a virtual object presents a missing portion of an ancient, damaged statue). In some examples, the virtual objects (122B, 124B, and / or 126B) may be displayed so as to occlude the corresponding real objects (122A, 124A, and / or 126A) (e.g., via active pixelated occlusion using a pixelated occlusion shutter). This may be desirable, for example, in instances where a virtual object acts as a visual replacement for a corresponding real object (such as in interactive storytelling applications where inanimate real objects become "living" characters).

[0034] In some examples, real objects (e.g., 122A, 124A, 126A) may be associated with virtual content or helper data that does not necessarily constitute a virtual object. The virtual content or helper data can facilitate processing or handling of the virtual object within a mixed reality environment. For example, such virtual content may include a two-dimensional representation of the corresponding real object, a custom asset type associated with the corresponding real object, or statistical data associated with the corresponding real object. This information can enable or facilitate calculations involving the real object without incurring unnecessary computational overhead.

[0035] In some embodiments, the presentation described above may also incorporate audio aspects. For example, in MRE 150, virtual monster 132 may be associated with one or more audio signals, such as footstep effects, that are generated as the monster walks around MRE 150. As described further below, a processor in mixed reality system 112 may calculate an audio signal corresponding to a mixed and processed combination of all such sounds within MRE 150 and present the audio signal to user 110 via one or more speakers included within mixed reality system 112 and / or one or more external speakers.

[0036] Exemplary Mixed Reality System

[0037] An exemplary mixed reality system 112 can include a wearable head device (e.g., a wearable augmented reality or mixed reality head device) comprising a display (which may be an eyepiece display, and may include left and right see-through displays and associated components for coupling light from the displays to the user's eyes), left and right speakers (e.g., positioned adjacent the user's left and right ears, respectively), an inertial measurement unit (IMU) (e.g., mounted on the temple arms of the head device), a quadrature coil electromagnetic receiver (e.g., mounted on the left temple component), left and right cameras (e.g., depth (time-of-flight) cameras) oriented away from the user, and left and right eye cameras oriented toward the user (e.g., to detect the user's eye movements). However, the mixed reality system 112 can incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic). Additionally, mixed reality system 112 may incorporate networking features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other mixed reality systems. Mixed reality system 112 may further include a battery (which may be mounted in an auxiliary unit, such as a belt pack designed to be worn around the user's waist), a processor, and memory. The wearable head device of mixed reality system 112 may include a tracking component, such as an IMU or other suitable sensor, configured to output a set of coordinates of the wearable head device relative to the user's environment. In some examples, the tracking component may provide input to a processor and implement simultaneous localization and mapping (SLAM) and / or visual odometry algorithms. In some examples, mixed reality system 112 may also include an auxiliary unit 320, which may be a handheld controller 300 and / or a wearable belt pack, as described further below.

[0038] 2A-2D illustrate components of an exemplary mixed reality system 200 (which may correspond to mixed reality system 112) that may be used to present an MRE (which may correspond to MRE 150) or other virtual environment to a user. FIG. 2A illustrates a perspective view of a wearable head device 2102 included within the exemplary mixed reality system 200. FIG. 2B illustrates a top view of the wearable head device 2102 worn on a user's head 2202. FIG. 2C illustrates a front view of the wearable head device 2102. FIG. 2D illustrates an edge view of an exemplary eyepiece 2110 of the wearable head device 2102. As shown in FIGS. 2A-2C, the exemplary wearable head device 2102 includes an exemplary left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an exemplary right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 can include a transmissive element through which the real environment is visible and a display element for presenting a display (e.g., via image-modulated light) that is overlaid on the real environment. In some embodiments, such display elements can include surface diffractive optical elements for controlling the flow of the image-modulated light. For example, the left eyepiece 2108 can include a left internal coupling grating set 2112, a left orthogonal pupil-extension (OPE) grating set 2120, and a left exit (output) pupil-extension (EPE) grating set 2122. Similarly, the right eyepiece 2110 can include a right internal coupling grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. The image-modulated light can be transferred to the user's eye via the internal coupling gratings 2112 and 2118, the OPEs 2114 and 2120, and the EPEs 2116 and 2122. Each internal coupling grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 can be designed to progressively deflect light downward toward its associated EPE 2122, 2116, thereby extending the exit pupil formed horizontally.Each EPE 2122, 2116 can be configured to progressively redirect at least a portion of the light received from its corresponding OPE grating set 2120, 2114 outward toward a user eyebox location (not shown), defined behind the eyepieces 2108, 2110, such that the exit pupil formed in the eyebox extends vertically. Alternatively, instead of the internal coupling grating sets 2112 and 2118, the OPE grating sets 2114 and 2120, and the EPE grating sets 2116 and 2122, the eyepieces 2108 and 2110 can include gratings and / or other arrangements of refractive and reflective features to control the coupling of image-modulated light into the user's eye.

[0039] In some examples, the wearable head device 2102 can include a left temple arm 2130 and a right temple arm 2132, where the left temple arm 2130 includes a left speaker 2134 and the right temple arm 2132 includes a right speaker 2136. A quadrature coil electromagnetic receiver 2138 can be located in the left temple assembly or another suitable location within the wearable head unit 2102. An inertial measurement unit (IMU) 2140 can be located in the right temple arm 2132 or another suitable location within the wearable head device 2102. The wearable head device 2102 can also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142, 2144 can preferably be oriented in different directions so that both cover a wider field of view.

[0040] 2A-2D , a left source of image-wise modulated light 2124 can be optically coupled into the left eyepiece 2108 through a left internal coupling grating set 2112, and a right source of image-wise modulated light 2126 can be optically coupled into the right eyepiece 2110 through a right internal coupling grating set 2118. The source of image-wise modulated light 2124, 2126 can include, for example, a fiber optic scanner, a projector including an electronic light modulator such as a digital light processing (DLP) chip or a liquid crystal on silicon (LCoS) modulator, or an emissive display such as a micro light emitting diode (μLED) or micro organic light emitting diode (μOLED) panel coupled into the internal coupling grating sets 2112, 2118 using one or more lenses per side. The input coupling grating sets 2112, 2118 can deflect light from the source of image-wise modulated light 2124, 2126 to an angle above the critical angle for total internal reflection (TIR) ​​for the eyepieces 2108, 2110. The OPE grating sets 2114, 2120 progressively deflect the propagating light downward by TIR towards the EPE grating sets 2116, 2122. The EPE grating sets 2116, 2122 progressively couple the light towards the user's face, including the pupils of the user's eyes.

[0041] In some embodiments, as shown in FIG. 2D , the left eyepiece 2108 and the right eyepiece 2110 each include multiple waveguides 2402. For example, each eyepiece 2108, 2110 can include multiple individual waveguides, each dedicated to a separate color channel (e.g., red, blue, and green). In some embodiments, each eyepiece 2108, 2110 can include multiple sets of such waveguides, each configured to impart a different wavefront curvature to the emitted light. The wavefront curvature may be convex with respect to the user's eye, for example, to present a virtual object positioned at a distance in front of the user (e.g., a distance corresponding to the inverse of the wavefront curvature). In some embodiments, the EPE grating sets 2116, 2122 can include curved grating grooves to impart a convex wavefront curvature by modifying the Poynting vector of light exiting across each EPE.

[0042] In some examples, stereoscopically adjusted left and right eye images can be presented to the user through light modulators 2124, 2126 and eyepieces 2108, 2110 for each image to create the perception that the displayed content is three-dimensional. The perceived realism of the presentation of three-dimensional virtual objects can be enhanced by selecting the waveguides (and thus the corresponding wavefront curvatures) so that the virtual objects are displayed at distances that approximate the distances indicated by the stereoscopic left and right images. This technique can also reduce motion sickness experienced by some users, which can be caused by differences between the depth perception cues provided by the stereoscopic left and right eye images and the automatic accommodation (e.g., object distance-dependent focus) of the human eye.

[0043] FIG. 2D illustrates an edge view from above of the right eyepiece 2110 of the exemplary wearable head device 2102. As shown in FIG. 2D , the plurality of waveguides 2402 can include a first subset of three waveguides 2404 and a second subset of three waveguides 2406. The two subsets of waveguides 2404, 2406 can be distinguished by different EPE gratings featuring different grating line curvatures to impart different wavefront curvatures to the exiting light. Within each of the subsets of waveguides 2404, 2406, each waveguide can be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although not shown in FIG. 2D , the structure of the left eyepiece 2108 is similar to that of the right eyepiece 2110.)

[0044] 3A illustrates example handheld controller components 300 of mixed reality system 200. In some examples, handheld controller 300 includes a grip portion 346 and one or more buttons 350 disposed along a top surface 348. In some examples, button 350 may be configured for use as an optical tracking target to track six degrees of freedom (6DOF) movement of handheld controller 300, for example, in conjunction with a camera or other optical sensor (which may be mounted in a head unit (e.g., wearable head device 2102) of mixed reality system 200). In some examples, handheld controller 300 includes a tracking component (e.g., an IMU or other suitable sensor) for detecting a position or orientation, such as a position or orientation relative to wearable head device 2102. In some examples, such a tracking component may be positioned in a handle of handheld controller 300 and / or may be mechanically coupled to the handheld controller. The handheld controller 300 can be configured to provide one or more output signals corresponding to one or more of a button press state, or the position, orientation, and / or movement (e.g., via an IMU) of the handheld controller 300. Such output signals may be used as inputs to a processor of the mixed reality system 200. Such inputs may correspond to the position, orientation, and / or movement of the handheld controller (or, for that matter, the position, orientation, and / or movement of a user's hand holding the controller). Such inputs may also correspond to a user pressing a button 350.

[0045] 3B illustrates an example auxiliary unit 320 of the mixed reality system 200. The auxiliary unit 320 can include a battery for providing energy to operate the system 200 and can include a processor for executing programs to operate the system 200. As shown, the example auxiliary unit 320 includes a clip 2128 for attaching the auxiliary unit 320 to a user's belt, etc. It will also be apparent that other form factors are suitable for the auxiliary unit 320, including form factors that do not involve mounting the unit on a user's belt. In some embodiments, the auxiliary unit 320 is coupled to the wearable head device 2102 through a multi-tube cable, which may include, for example, electrical wires and optical fibers. A wireless connection between the auxiliary unit 320 and the wearable head device 2102 can also be used.

[0046] In some examples, mixed reality system 200 can include one or more microphones to detect sound and provide a corresponding signal to the mixed reality system. In some examples, the microphones may be attached to or integrated with wearable head device 2102 and configured to detect the user's voice. In some examples, microphones may be attached to or integrated with handheld controller 300 and / or auxiliary unit 320. Such microphones may be configured to detect environmental sounds, ambient noise, the user's or a third party's voice, or other sounds.

[0047] 4 shows an example functional block diagram that may correspond to an example mixed reality system, such as mixed reality system 200 described above (which may correspond to mixed reality system 112 with respect to FIG. 1). As shown in FIG. 4, example handheld controller 400B (which may correspond to handheld controller 300 (“totem”)) includes a totem / wearable head device six degrees of freedom (6DOF) totem subsystem 404A, and example wearable head device 400A (which may correspond to wearable head device 2102) includes a totem / wearable head device 6DOF subsystem 404B. In an example, 6DOF totem subsystem 404A and 6DOF subsystem 404B cooperate to determine six coordinates of handheld controller 400B relative to wearable head device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom may be expressed relative to the coordinate system of wearable head device 400A. The three translational offsets may be represented as X, Y, and Z offsets within such a coordinate system, a translation matrix, or some other representation. The rotational degrees of freedom may be represented as a sequence of yaw, pitch, and roll rotations, as a rotation matrix, as a quaternion, or some other representation. In some examples, the wearable head device 400A, one or more depth cameras 444 (and / or one or more non-depth cameras) included within the wearable head device 400A, and / or one or more optical targets (e.g., buttons 350 of handheld controller 400B as described above or dedicated optical targets included within handheld controller 400B) can be used for 6DOF tracking. In some examples, the handheld controller 400B can include a camera as described above, and the wearable head device 400A can include an optical target for optical tracking in conjunction with the camera. In some embodiments, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids, which are used to wirelessly transmit and receive three distinguishable signals.By measuring the relative magnitudes of the three distinguishable signals received in each of the coils used to receive, the 6DOF of the wearable head device 400A relative to the handheld controller 400B can be determined. Additionally, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU), which is useful for providing improved accuracy and / or more timely information regarding high speed movements of the handheld controller 400B.

[0048] In some examples, it may be necessary to transform coordinates from a local coordinate space (e.g., a coordinate space that is fixed relative to the wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space that is fixed relative to the real environment), e.g., to compensate for movement of the wearable head device 400A relative to coordinate system 108. For example, such a transformation may be necessary so that the display of the wearable head device 400A presents virtual objects in an expected position and orientation relative to the real environment (e.g., a virtual person sitting in a real chair facing forward, regardless of the position and orientation of the wearable head device), rather than in a fixed position and orientation on the display (e.g., the same position in the bottom right corner of the display), preserving the illusion that the virtual objects exist in the real environment (and do not appear unnaturally positioned in the real environment, e.g., as the wearable head device 400A shifts and rotates). In some examples, a compensatory transformation between coordinate spaces can be determined by processing images from depth camera 444 using SLAM and / or visual odometry procedures to determine the transformation of wearable head device 400A relative to coordinate system 108. In the example shown in FIG. 4 , depth camera 444 is coupled to SLAM / visual odometry block 406 and can provide images to block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the images and then determine the position and orientation of the user's head, which can be used to identify a transformation between the head coordinate space and another coordinate space (e.g., an inertial coordinate space). Similarly, in some examples, an additional source of information about the user's head pose and location is obtained from IMU 409. Information from IMU 409 can be integrated with information from SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information for rapid adjustments of the user's head pose and position.

[0049] In some examples, depth camera 444 can provide 3D images to hand gesture tracker 411, which can be implemented within a processor of wearable head device 400A. Hand gesture tracker 411 can identify the user's hand gestures, for example, by matching the 3D images received from depth camera 444 to stored patterns representing hand gestures. Other suitable techniques for identifying the user's hand gestures will also be apparent.

[0050] In some embodiments, one or more processors 416 may be configured to receive data from the wearable head device's 6DOF headgear subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, and / or hand gesture tracker 411. The processor 416 may also send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A, such as in embodiments in which the handheld controller 400B is untethered. The processor 416 may further communicate with additional components, such as an audio / visual content memory 418, a graphical processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to a head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to a left source of imagewise modulated light 424 and a right channel output coupled to a right source of imagewise modulated light 426. The GPU 420 may output stereoscopic image data to the sources of imagewise modulated light 424, 426, for example, as described above with respect to FIGS. 2A-2D . The DSP audio spatializer 422 may output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 may receive an input from the processor 419 indicating a direction vector from the user to a virtual sound source (which may be moved by the user, e.g., via the handheld controller 320). Based on the direction vector, the DSP audio spatializer 422 may determine a corresponding HRTF (e.g., by accessing an HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the believability and realism of virtual sounds by incorporating the user's relative position and orientation to the virtual sounds in the mixed reality environment, i.e., by presenting virtual sounds that match the user's expectations of what they would hear if the virtual sounds were real sounds in a real environment.

[0051] 4 , one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 may be included in auxiliary unit 400C (which may correspond to auxiliary unit 320 described above). Auxiliary unit 400C may include battery 427 to power its components and / or provide power to wearable head device 400A or handheld controller 400B. Including such components in an auxiliary unit, which may be mounted on the user's waist, can limit the size and weight of wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0052] While Figure 4 presents elements corresponding to various components of an exemplary mixed reality system, various other suitable arrangements of these components will be apparent to those skilled in the art. For example, elements shown in Figure 4 as associated with auxiliary unit 400C may instead be associated with wearable head device 400A or handheld controller 400B. Furthermore, some mixed reality systems may dispense with handheld controller 400B or auxiliary unit 400C entirely. Such variations and modifications should be understood as being within the scope of the disclosed embodiments.

[0053] Mixed Reality Video Capture

[0054] Capturing and recording an AR / MR experience can be more complex than capturing two-dimensional video and may present new challenges. For example, a two-dimensional video recording system may capture only one stream of information (e.g., what an optical / image sensor on the video recording system can capture). On the other hand, an AR / MR experience may have two or more streams of information. For example, one stream of information may be two-dimensional video, and another stream of information may be rendered virtual content. Sometimes, three or more streams of information may be used (e.g., when stereo video is captured from two two-dimensional videos along with accompanying virtual content). Combining multiple streams of information into an accurate AR / MR recording may include considering differences in line of sight between the multiple streams. Furthermore, it may be beneficial to handle the combination process in a computationally efficient manner, especially for portable AR / MR systems that may have limited battery and / or processing power.

[0055] A user may both capture and record AR / MR content. In some embodiments, a user may want to record and share personal AR / MR content on social media (e.g., an AR / MR recording of their children playing). In some embodiments, a user may want to record and share commercial AR / MR content (e.g., an advertisement). The recorded AR / MR content may be shared both to a two-dimensional screen (e.g., by combining one stream of virtual content with one stream of real content) and to other MR systems (e.g., by combining two streams of virtual content with two streams of real content). While the original user of the AR / MR system may experience the real environment through transmissive optics (e.g., through a partially transmissive lens), the viewer viewing the recorded AR / MR content may not be present in the same real environment. Presenting the complete AR / MR recording to the viewer may therefore include a video capture of the real environment so that the virtual content can be overlaid on the video capture of the real environment.

[0056] Capturing an MRE in a video recording may present additional challenges in addition to displaying the MRE. For example, mixed reality video capture may include two recording streams: a real environment recording stream and a virtual environment recording stream. The real environment recording stream can be captured using one or more optical sensors (e.g., cameras). It may be advantageous to mount one or more optical sensors on the AR / MR system (e.g., MR system 112, 200). In embodiments in which the AR / MR system is a wearable head device, one or more optical sensors can be mounted on the wearable head device in close proximity to, and optionally facing in the same direction as, the user's eyes. In the described mounted system, the one or more optical sensors can be positioned to capture video recording of the real environment in a manner that closely approximates how a user experiences the real environment.

[0057] The virtual environment recording stream can be captured by recording the virtual content being displayed to the user. In some embodiments, the AR / MR system (e.g., MR system 112, 200) can simultaneously display two streams of virtual content to the user (e.g., one stream per eye), providing the user with a stereoscopic view. Each virtual rendering can be rendered from a slightly different eye line (which may correspond to a difference in eye line between the user's left and right eyes), and the difference in eye line can make the virtual content appear three-dimensional. The virtual environment recording stream can include one or both of the virtual renderings displayed to the user's eyes. In some embodiments, the virtual rendering can be overlaid with the real environment recording stream to simulate the MRE the user experiences when using the AR / MR system. It can be beneficial to use virtual renderings rendered for the user's eyes that are physically closest in location to one or more optical sensors used for the real environment recording stream. The physical proximity of the user's eyes (and the virtual renderings rendered for their corresponding eye lines) to the optical sensors can make the combination process more accurate due to minimal shifts in eye line.

[0058] One problem that may arise from combining a real environment recording stream and a virtual environment recording stream is that the two recording streams may not have the same resolution and / or aspect ratio. For example, the real environment recording stream may capture video at a resolution of 1,920 x 1,080 pixels, while the virtual environment recording stream may render virtual content only at a resolution of 1,280 x 720 pixels. Another problem that may arise is that the real environment recording stream may not have the same field of view as the virtual environment recording stream. For example, a camera may capture the real environment recording stream using a larger field of view than the rendered virtual environment recording stream. Thus, more information (e.g., a larger field of view) may be captured in the real environment recording stream than in the virtual environment recording stream. In some embodiments, real content that is outside the field of view of the virtual content may not have corresponding rendered virtual content in the virtual rendering. When a real environment recording stream is merged with a virtual environment recording stream, an artificial "border" may exist in the combined video, outside of which the virtual content is not displayed, but the real content is still recorded and displayed. This border may be annoying to the viewer because it may distract from the integration of the virtual and real content. This may further present inaccuracies to the viewer. For example, a real object recorded by a camera may have corresponding virtual content (e.g., an information overlay). However, if the real object is outside the field of view of the virtual rendering but still within the field of view of the camera, the corresponding virtual content may not be displayed to the viewer until the real object moves within the artificial border. This behavior may confuse viewers of the AR / MR recording because the corresponding virtual content of the real object may or may not be visible depending on where the real object is located relative to the border.

[0059] 5A illustrates an example embodiment that may reduce and / or eliminate the boundary between virtual and real content. The example embodiment shown can be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above. In the depicted embodiment, an RGB camera can be positioned near the user's left eye. The RGB camera can record view 502, which can overlap with and / or include view 504 (and / or view 506) of a virtual rendering that is presented to the user's left eye. View 502 may be larger than view 504, and therefore combining (e.g., overlaying or superimposing) view 504 onto view 502 may result in an artificial boundary 503. In this embodiment, the AR / MR recording may display real content in the entirety of view 502, but virtual content may only be shown within boundary 503. Boundary 503 can be annoying to the viewer because virtual content may exist outside of boundary 503 (at least within view 502), but the virtual content may not be visible until the virtual content moves within boundary 503. One solution to reduce and / or eliminate boundary 503 is to crop view 502 to the size and location of boundary 503. The cropped view may result in the virtual content being sized in the same manner as the real content, creating an integrated and immersive viewing experience. However, the cropped view may reduce the viewable area for the viewer by "discarding" information captured by the RGB camera.

[0060] 5B illustrates an example embodiment for merging the lines of sight from view 502 and view 504. The example embodiment shown can be implemented using one or more components of a mixed reality system, such as one or more of wearable head device 2102, handheld controller 300, and auxiliary unit 320 of example mixed reality system 200 described above. Because RGB camera 508 may have a focal point 508 that is offset from focal point 510 (which may correspond to the center of the user's left eye), view 502 may present a different line of sight than view 504, resulting in virtual content that may not be “synchronized” (e.g., aligned and / or matched) with the real content. This can be problematic because the virtual content may not correspond properly to the real content (e.g., if a virtual object's face is visible from the line of sight of the user's eyes but should not be visible from the line of sight of the RGB camera), which may destroy the immersive feel of the AR / MR recording. One solution may be to "map" view 504 to view 502 using planar projection. In planar projection, one or more points P1 in view 504 may be mapped onto a line between the focal point 508 and one or more corresponding points P2 in view 502 (or vice versa). This may be done using a computer vision algorithm, which may determine that an observation at one or more points P1 is an observation of the same features as an observation at one or more points P2. The mapped combination may mitigate differences in line of sight due to different focal points 508, 510. In some embodiments, the combined AR / MR recording may display the virtual content in approximately the same way as it would appear to a user of the AR / MR system. However, as described above, in some embodiments, useful information captured by an RGB camera may be discarded to avoid an artificial boundary between the real and virtual content.

[0061] FIG. 6 illustrates an example embodiment for producing an AR / MR recording. The example embodiment shown can be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above. The AR / MR system (e.g., MR system 112, 200) can render the virtual content from one or more eyelines (e.g., the virtual content can be rendered once from the eyeline of the user's left eye and once from the eyeline of the user's right eye). The AR / MR system also renders the virtual content a third time from the eyeline of an RGB camera located near the user's eye (e.g., the user's left eye) and which may face approximately the same direction as the user's forward line of sight. In some embodiments, this third pass rendering of the virtual content can be rendered to the same resolution, aspect ratio, and / or field of view of the RGB camera. This can enable AR / MR recordings to be generated with synchronized virtual and real content without discarding any information from either the RGB camera or the virtual rendering. However, in some embodiments, third-pass rendering of the virtual content can be computationally expensive. The third-pass rendering can include a full geometry pass, which can include full draw calls and / or rasterization. Additional details can be found in U.S. Patent Application No. 15 / 924,144, the contents of which are incorporated herein by reference in their entirety. Performing third-pass rendering on a portable AR / MR system with limited computing resources can be infeasible (e.g., because the AR / MR system already renders twice for the user's two eyes). Furthermore, the higher computational load can result in additional power consumption, which can have a negative impact on battery life for AR / MR systems that rely on portable power.

[0062] 7 illustrates an example embodiment for producing an AR / MR recording. The example embodiment shown can be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above. The AR / MR system (e.g., MR system 112, 200) can render virtual content from one or more eyelines (e.g., virtual content can be rendered once from the eyeline of the user's left eye and once from the eyeline of the user's right eye). In some embodiments, the two virtual renderings (e.g., one per eye) can be non-uniform. For example, the virtual rendering for the user's left eye may render for view 702, which may be different from the virtual rendering for the user's right eye, which may render for view 706. View 706 may be approximately the same size as the field of view for the user's right eye (e.g., therefore, virtual content that would not be displayed to the user is not rendered, thereby minimizing computational load). View 702 may be larger than view 706 and may render a larger field of view than that visible to the user's left eye. View 702 may include view 704, which may be the field of view available to the user's left eye (e.g., all virtual content displayed to the user may be included within view 702, in addition to some virtual content that may not be displayed to the user). View 702 may optionally have the same or approximately the same field of view as view 708, which may be the field of view available to an RGB camera (e.g., an RGB camera placed near the user's left eye and facing the same direction as the user's forward gaze).

[0063] During AR / MR recording, views 702 and 706 can be rendered (e.g., by MR system 112, 200). View 702 can be cropped to view 704, which can have the same field of view as view 706, but view 704 can be rendered from a slightly different perspective (e.g., from the left eye's perspective instead of the right eye's perspective). Views 704 and 706 can display virtual content to the user, and views 704 and 706 can create stereoscopic images and simulate three-dimensional virtual content. An RGB camera can simultaneously record video of the real environment from a perspective similar to the user's perspective (e.g., the RGB camera may be placed near one of the user's eyes). This real environment recording stream can then be combined with view 702 (which can be a virtual environment recording stream) to create an MR recording. In some embodiments, the virtual environment recording stream from view 702 can be approximately the same resolution, aspect ratio, and / or field of view as view 708. Thus, AR / MR recording, which combines real-world and virtual-world recording streams, may discard little or none of the information captured by the RGB camera and / or virtual rendering. Furthermore, it may be beneficial to render non-uniform stereoscopic virtual content because extending one of the two rendering passes may be more computationally efficient than using a full third-pass rendering. Extending an existing rendering pass may only involve additional rasterization for additional pixels and may not involve a full geometry pass (unlike third-pass rendering in some embodiments).

[0064] 8 illustrates an example process for generating an AR / MR recording, which may be an augmented video capture. The example process shown can be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above. In step 802, a request to start augmented video capture may be received. The request can be user-initiated (e.g., the user selects a setting to start recording, or the user speaks aloud to start recording, which the AR / MR system may process as a user request, or the user performs a gesture, which the AR / MR system may process as a user request) or automated (e.g., a computer vision algorithm detects a scene that a machine learning algorithm determines is likely to be of interest to the user, and the MR system therefore automatically starts recording). In step 804, the user's pose and the RGB camera pose may be determined / estimated. The user's pose may include information about the user's position and / or orientation within the three-dimensional environment. The user's pose may be the same as or similar to the pose of the AR / MR system (e.g., if the AR / MR system is a wearable head device that may be fixed relative to the user's head). The pose may be estimated using methods such as SLAM, visual inertial odometry, and / or other suitable methods. The RGB camera's pose may be determined / estimated independently and / or derived from the AR / MR system's pose (which may approximate the user's pose). For example, the RGB camera may be mounted on the AR / MR system at a fixed location that may be known to the AR / MR system. The AR / MR system may utilize its own estimated pose in combination with the RGB camera's location relative to the AR / MR system to determine / estimate the RGB camera's pose.In step 806, virtual content can be rendered for a first eye of the user (e.g., the user's right eye) based on pose estimates for the user and / or the AR / MR system. The virtual content is rendered by simulating a first virtual camera that can be positioned at the center of the user's eye (e.g., the user's first eye).

[0065] In step 808, virtual content may be rendered for the user's second eye (e.g., the user's left eye) based on pose estimates for the user and / or the AR / MR system and the RGB camera. The virtual content rendered for the user's second eye may simulate a second virtual camera that may be positioned at the center of the user's eye (e.g., the user's second and / or left eye). The virtual content rendered for the user's second eye may include changes in line of sight and / or virtual camera location as well as other changes compared to the virtual content rendered for the user's first eye. In some embodiments, the frame buffer size (e.g., image plane size) of the second virtual camera (e.g., a virtual camera to be composited with the RGB camera, which may also be a virtual camera for the user's left eye) may be increased as compared to the frame buffer size of the virtual camera for the user's first eye (e.g., the first virtual camera). The aspect ratio of the frame buffer for the second eye may match the aspect ratio of the frame buffer for the first eye (although other aspect ratios may be used). In some embodiments, the projection matrix of the virtual camera for the second eye may be changed so that a larger field of view is rendered compared to the virtual camera for the first eye. It may be desirable to change the virtual camera's projection matrix to account for the larger frame buffer size. Failure to increase the field of view and / or change the virtual camera's projection matrix to account for the larger frame buffer size may result in virtual content that does not match the virtual content presented to the different eye, such that the stereoscopic effect may be reduced and / or lost. In some embodiments, the increased field of view and / or changed projection matrix may produce a virtual rendering with the same field of view as an RGB camera located near the user's second eye. In some embodiments, the view matrix may remain unchanged so that the virtual camera is composited with the RGB camera (e.g., the virtual camera may still render from the same position with the same vector relative to the target).In some embodiments, a minimum clipping distance (e.g., the minimum distance from a user's eye to virtual content below which the virtual content cannot be rendered) may prevent virtual objects from being rendered in regions that are outside the field of view of the RGB camera but may be inside the field of view of a virtual camera centered on the user's second eye.

[0066] In step 810, the virtual content rendered for the second eye in step 808 may be cropped for display to the user's second eye. In some embodiments, the center of the expanded field of view can also be the center of the cropped field of view presented to the user's second eye. The cropped field of view can be the same size as the field of view for the virtual content rendered for the user's first eye. The virtual content rendered for the user's first eye and the cropped virtual content rendered for the user's second eye can be combined to produce a stereoscopic effect for the virtual content displayed to the user, which can simulate three-dimensionality of the virtual content. In step 812, the virtual content may be displayed to the user. In some embodiments, steps 806, 808, and / or 810 can occur simultaneously or substantially simultaneously. For example, an AR / MR system may render virtual content for a user's first eye and the user's second eye, crop the rendered virtual content for the user's second eye, and display stereoscopic virtual content to the user with little or no perceptible delay (e.g., the stereoscopic virtual content may track the real content as the user looks around the real environment).

[0067] At step 814, RGB video capture may begin (e.g., using an RGB camera mounted on the AR / MR system). At step 816, an augmented (e.g., uncropped) virtual content field of view may be composited with the captured RGB video. For example, the augmented virtual content field of view may be projected onto the captured RGB video (e.g., using planar projection). In some embodiments, the captured RGB video may be projected onto the virtual content field of view (e.g., using planar projection). In some embodiments, the virtual content field of view may be the same size as the RGB camera's field of view. This may result in very little loss of RGB camera information and / or virtual information. In some embodiments, some information may be lost as a result of the projection process. In some embodiments, the virtual content field of view may be larger than the RGB camera's field of view, and therefore, losses due to projection may not require discarding information captured by the RGB camera. In some embodiments, the virtual content field of view may be smaller than the RGB camera's field of view, and therefore, losses due to projection may require discarding information captured by the RGB camera. At step 818, the augmented video capture may be configured and stored. In some embodiments, the augmented video capture may display more information than was presented to the user of the recording AR / MR system. For example, the AR / MR recording may have an augmented field of view compared to the user's field of view while using the AR / MR system.

[0068] 9 illustrates an example system for configuring and storing augmented video capture, which may correspond to steps 816 and 818. The example system shown can be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above. Camera image frames from camera 902 can be sent to a compositor 906, along with pose data for the user's head / eyes and / or pose data for the camera from IMU 904. The pose data can then be sent from the compositor 906 to an AR / MR application and / or a graphics processing unit (“GPU”) 908, which can pull information (e.g., models) from a 3D database. The AR / MR application 908 can send an augmented virtual rendering to the compositor 906. The compositor 906 can then compose a virtual rendering (e.g., using planar projection) augmented with the camera image frames from the camera 902 and send the composed frames to the media encoder 912. The media encoder 912 can then send the encoded image frames to the recording database 914 for storage.

[0069] While an embodiment involving one RGB camera is described above, it is also contemplated that the systems and methods described herein can be applied to any number of RGB cameras. For example, two RGB cameras may be used, with one RGB camera mounted near each of the user's eyes. The virtual rendering for each eye may then be augmented (although only a limited field of view may be presented to the user during use) and projected onto the field of view of the RGB camera (the RGB camera image may also be projected into the virtual image). This may enable stereoscopic augmented video capture, which may be particularly suitable for playback on another MR system, which may provide a more immersive experience than playback on a 2D screen. It is also contemplated that the systems and methods may be used to shift the line of sight of the RGB camera as close as possible to that of the user's eyes. For example, mirrors and / or other optical elements may be used to reflect light so that the RGB camera views the content from the same (or nearly the same) line of sight as the user's eyes. The augmented video capture may also include synthesized audio in addition to the synthesized video. For example, a virtual audio signal may be recorded and combined with a recorded real audio signal (which may be captured by one or more microphones on the MR system). The combined audio may also be synchronized with the combined video.

[0070] Although the disclosed embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are to be understood as being included within the scope of the disclosed embodiments as defined by the appended claims.

Claims

1. 1. A method, comprising: receiving an image of a real environment, the image having a first field of view; receiving an image of a virtual environment, the image having a second field of view different from the first field of view; generating a combined image based on the image of the virtual environment and the image of the real environment; presenting the combined image to a first eye of a user; and Including, generating the combined image includes performing a planar projection, the planar projection including mapping points of the image of the virtual environment to corresponding points of the image of the real environment; 11. The method of claim 10, wherein the image of the virtual environment is rendered using a first frame buffer having a first size, the first size being larger than a second size of a second frame buffer, the second frame buffer being used to render a second image of the virtual environment configured to be presented to a second eye of the user.

2. The method of claim 1, wherein the image of the virtual environment is rendered using a minimum clipping distance, the minimum clipping distance capable of preventing rendering of one or more virtual objects in an area outside the first field of view but inside the second field of view.

3. The method comprises: generating a cropped image based on the image of the virtual environment; presenting a stereoscopic image based on the cropped image; The method of claim 1 further comprising:

4. The method of claim 3 , wherein the cropped image has a field of view that is smaller than a field of view of the image of the virtual environment.

5. The method described in claim 3, wherein the cropped image is rendered using a third frame buffer having a size identical to the first size of the first frame buffer used for the image of the virtual environment.

6. The method of claim 1 , wherein the image of the virtual environment has an aspect ratio and the image of the real environment has an aspect ratio that is the same as the aspect ratio of the image of the virtual environment.

7. The method of claim 1 , wherein the image of the real environment has a first resolution and the image of the virtual environment has a second resolution different from the first resolution.

8. The method of claim 1 , wherein receiving the image of the real environment comprises receiving the image via a camera.

9. The method of claim 8 , wherein the camera is configured to be placed to the left of the user's left eye.

10. The method of claim 8 , wherein the camera is configured to be placed to the right of the user's right eye.

11. The method of claim 8 , wherein the camera is configured to be placed between the user's left eye and the user's right eye.

12. generating the combined image includes generating the combined image via a first device; The method of claim 1 , wherein presenting the combined image includes presenting the combined image via a display of a second device different from the first device.

13. generating the combined image includes generating the combined image via a first device; The method of claim 1 , wherein presenting the combined image comprises presenting the combined image via a display of the first device.

14. The method of claim 1 , wherein the first field of view is larger than the second field of view.

15. 1. A system comprising: The system comprises one or more processors configured to perform a method; The method comprises: receiving an image of a real environment, the image having a first field of view; receiving an image of a virtual environment, the image having a second field of view different from the first field of view; generating a combined image based on the image of the virtual environment and the image of the real environment; presenting the combined image to a first eye of a user; and Including, generating the combined image includes performing a planar projection, the planar projection including mapping points of the image of the virtual environment to corresponding points of the image of the real environment; The system, wherein the image of the virtual environment is rendered using a first frame buffer having a first size, the first size being larger than a second size of a second frame buffer, the second frame buffer being used to render a second image of the virtual environment configured to be presented to a second eye of the user.

16. The system described in claim 15, wherein the image of the virtual environment is rendered using a minimum clipping distance, the minimum clipping distance capable of preventing rendering of one or more virtual objects in an area outside the first field of view but inside the second field of view.

17. generating the combined image includes generating the combined image via a first device; The system of claim 15 , wherein presenting the combined image includes presenting the combined image via a display of a second device different from the first device.

18. generating the combined image includes generating the combined image via a first device; The system of claim 15 , wherein presenting the combined image includes presenting the combined image via a display of the first device.

19. The system of claim 15 , wherein the first field of view is larger than the second field of view.

20. A non-transitory computer-readable medium storing instructions, comprising: The instructions, when executed by one or more processors, receiving an image of a real environment, the image having a first field of view; receiving an image of a virtual environment, the image having a second field of view different from the first field of view; generating a combined image based on the image of the virtual environment and the image of the real environment; presenting the combined image to a first eye of a user; and causing the one or more processors to perform a method comprising: generating the combined image includes performing a planar projection, the planar projection including mapping points of the image of the virtual environment to corresponding points of the image of the real environment; The image of the virtual environment is rendered using a first frame buffer having a first size, the first size being larger than a second size of a second frame buffer, the second frame buffer being used to render a second image of the virtual environment configured to be presented to a second eye of the user.

21. The non-transitory computer-readable medium of claim 20, wherein the image of the virtual environment is rendered using a minimum clipping distance that is capable of preventing rendering of one or more virtual objects in an area that is outside the first field of view but inside the second field of view.

22. generating the combined image includes generating the combined image via a first device; 21. The non-transitory computer-readable medium of claim 20, wherein presenting the combined image comprises presenting the combined image via a display of the first device.

23. 21. The non-transitory computer-readable medium of claim 20, wherein the first field of view is larger than the second field of view.

Citation Information

Patent Citations

  • Information processing method, and information processing device

    JP2007299062A

  • Display control program, display control device, display control system and display control method

    JP2012104023A

  • Information display system, information display method and program for information display

    JP2013054661A

  • Hybrid reality for 3D human-machine interfaces

    JP2014505917A

  • Imaging device, head-mounted display, information processing system, and information processing method

    JP2017204674A