Reverse pass-through glasses for augmented and virtual reality devices
The integration of a near-eye display with a light field display and eye imaging system in AR and VR devices addresses the lack of realistic facial feature viewing, enabling high-resolution, three-dimensional projections for improved social interactions.
Patent Information
- Application Number
- JP2023538972
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-17
- Filing Date
- 2021-12-23
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Current AR and VR devices lack the ability to provide a realistic, three-dimensional view of a user's facial features to onlookers, leading to a disconnect between the user and their surroundings, which can be frustrating or dangerous in social interactions.
The implementation of a near-eye display with a light field display and eye imaging system that captures and projects a three-dimensional model of the user's face to onlookers, using a multi-lenslet array and pixel array to create a high-resolution, autostereoscopic image that adapts to the viewer's perspective.
This solution enhances face-to-face interactions by providing a realistic, high-resolution, three-dimensional view of the user's face, allowing for more natural and engaging social interactions, bridging the gap between VR and AR experiences.
Smart Images

Figure 0007825624000013 
Figure 0007825624000014 
Figure 0007825624000015
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to augmented reality (AR) and virtual reality (VR) devices that include a reverse pass-through feature that provides a realistic view of a user's facial features to a front onlooker. More specifically, the present disclosure provides an autostereoscopic external display for an AR / VR headset user's onlooker. [Background technology]
[0002] In the field of AR and VR devices, some devices include outward-facing displays that provide an onlooker with an image that is being displayed to the device user. While these configurations can help onlookers better understand what the AR or VR device user is experiencing, they can leave onlookers unaware of the user's state of mind or focus of attention, such as when the user is using a pass-through mode to communicate with the onlooker and is not otherwise participating in the virtual reality environment. Furthermore, for such devices with outward-facing displays, the devices are typically conventional two-dimensional displays that lack a realistic view of a full-body image of at least a portion of the user's face or head, such as depicting the accurate depth and distance of the user's face or head within the device. Summary of the Invention
[0003] According to a first aspect of the present disclosure, there is provided a device comprising a near-eye display configured to provide an image to an object, an eye imaging system configured to collect images of the object, and a light field display configured to provide an autostereoscopic image of a three-dimensional model of the object to an onlooker, wherein the autostereoscopic image includes perspective-corrected views of the object from multiple viewpoints within a field of view of the light field display.
[0004] In some embodiments, the light field display comprises a pixel array and a multi-lenslet array, the pixel array configured to provide a segmented view of an object to the multi-lenslet array, the segmented view comprising multiple portions of the field of view of the light field display at a selected viewpoint.
[0005] In some embodiments, the eye imaging system comprises two cameras to collect a binocular view of the subject.
[0006] In some embodiments, the device further comprises one or more processors and a memory that stores instructions that, when executed by the one or more processors, generate a three-dimensional representation of the object from images of the object.
[0007] In some embodiments, the near-eye display provides a subject with a three-dimensional representation of the environment, including the spectators.
[0008] In some embodiments, the eye imaging system includes an infrared camera that receives images from the subject in a reflective mode from a dichroic mirror adjacent to the light field display.
[0009] In some embodiments, the light field display comprises a microlens array having a plurality of microlenses arranged in a two-dimensional pattern with a preselected pitch to avoid crosstalk between perspective-corrected views for two viewpoints for the viewer.
[0010] In some embodiments, the light field display further comprises an immersion stop adjacent to the microlens array, the immersion stop including a plurality of openings such that each opening is aligned with the center of each microlens in the microlens array.
[0011] In some embodiments, the light field display includes a pixel array divided into a plurality of active segments, each active segment in the pixel array having a dimension corresponding to the diameter of a refractive element in the multi-lenslet array.
[0012] In some embodiments, the device further comprises one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the light field display to divide the pixel array into a plurality of active segments, each active segment configured to provide a portion of the field of view of the light field display at a selected viewpoint for an onlooker.
[0013] According to a second aspect of the present disclosure, a computer-implemented method is provided that includes receiving a plurality of images from one or more headset cameras having at least two or more views of a subject, the subject being a headset user; extracting a plurality of image features from the images using a set of learnable weights; forming a three-dimensional model of the subject using the set of learnable weights; mapping the three-dimensional model of the subject to an autostereoscopic display format that associates an image projection of the subject with a selected observation point for a spectator; and providing the image projection of the subject on a display of a device when the spectator is located at the selected observation point.
[0014] In some embodiments, extracting image features includes extracting unique characteristics of the headset camera used to collect each image.
[0015] In some embodiments, mapping the three-dimensional model of the object to an autostereoscopic display format includes interpolating a feature map associated with a first observation point with a feature map associated with a second observation point.
[0016] In some embodiments, mapping the three-dimensional model of the object to an autostereoscopic display format includes aggregating image features of a plurality of pixels along the direction of the selected observation point.
[0017] In some embodiments, mapping the three-dimensional model of the object to an autostereoscopic display format includes concatenating multiple feature maps generated by each headset camera in a permutation-invariant combination, each headset camera having unique characteristics.
[0018] In some embodiments, providing an image projection of the object includes providing a second image projection on the device display as the spectator moves from the first observation point to the second observation point.
[0019] According to a third aspect of the present disclosure, there is provided a computer-implemented method for training a model to provide views of an object on an autostereoscopic display in a virtual reality headset, the method including: collecting a plurality of ground truth images of a plurality of users' faces; correcting the ground truth images with stored and calibrated pairs of stereoscopic images; mapping a three-dimensional model of the object to an autostereoscopic display format that associates an image projection of the object with a selected observation point for an onlooker; determining a loss value based on a difference between the ground truth images and the image projection of the object; and updating the three-dimensional model of the object based on the loss value.
[0020] In some embodiments, generating the multiple synthetic views includes projecting image features from each of the ground truth images along a selected observation direction and concatenating the multiple feature maps generated by each ground truth image in a permutation-invariant combination, where each ground truth image has unique characteristics.
[0021] In some embodiments, training the three-dimensional model of the object includes updating at least one of the set of learnable weights for each of the plurality of features based on a value of a loss function indicative of a difference between the ground truth image and the image projection of the object.
[0022] In some embodiments, training the three-dimensional model of the object includes training a background value for each of a plurality of pixels in the ground truth image based on pixel background values projected from a plurality of ground truth images.
[0023] According to a fourth aspect of the present disclosure, there is provided a system including first means for storing instructions and second means for executing the instructions to implement a method, the method including receiving a plurality of two-dimensional images having at least two or more views of an object; extracting a plurality of image features from the two-dimensional images using a set of learnable weights; projecting the image features along a direction between a three-dimensional model of the object and a selected observation point for a spectator; and providing an autostereoscopic image of the three-dimensional model of the object to the spectator.
[0024] It will be understood that any feature described herein as suitable for incorporation into one or more aspects or embodiments of the present disclosure is intended to be generalizable across any and all aspects and embodiments of the present disclosure. Other aspects of the present disclosure can be appreciated by those skilled in the art in light of the detailed description, claims, and drawings of the present disclosure. The foregoing general description and the following detailed description are exemplary and explanatory only and do not limit the scope of the claims. [Brief explanation of the drawings]
[0025] [Figure 1A] 1 illustrates an AR or VR device including an autostereoscopic external display, according to some embodiments. [Figure 1B] 1 illustrates a user of an AR or VR device as seen by a front onlooker, according to some embodiments. [Figure 2]FIG. 10 is a detailed view of an eyepiece for an AR or VR device configured to provide a forward onlooker with a reverse pass-through view of a user's face, according to some embodiments. [Figure 3A] 1 illustrates different aspects and components of a microlens array used to provide a forward onlooker with a reverse pass-through view of an AR or VR device user, according to some embodiments. [Figure 3B] 1 illustrates different aspects and components of a microlens array used to provide a forward onlooker with a reverse pass-through view of an AR or VR device user, according to some embodiments. [Figure 3C] 1 illustrates different aspects and components of a microlens array used to provide a forward onlooker with a reverse pass-through view of an AR or VR device user, according to some embodiments. [Figure 3D] 1 illustrates different aspects and components of a microlens array used to provide a forward onlooker with a reverse pass-through view of an AR or VR device user, according to some embodiments. [Figure 4] 1 illustrates a ray tracing view through a light field display to provide a front onlooker with a wide-angle, high-resolution view of an AR or VR device user, according to some embodiments. [Figure 5A] 1 illustrates different aspects of resolution characteristics in a microlens array used to provide a wide-angle, high-resolution view for a user of an AR or VR device, according to some embodiments. [Figure 5B] 1 illustrates different aspects of resolution characteristics in a microlens array used to provide a wide-angle, high-resolution view for a user of an AR or VR device, according to some embodiments. [Figure 5C] 1 illustrates different aspects of resolution characteristics in a microlens array used to provide a wide-angle, high-resolution view for a user of an AR or VR device, according to some embodiments. [Figure 5D]1 illustrates different aspects of resolution characteristics in a microlens array used to provide a wide-angle, high-resolution view for a user of an AR or VR device, according to some embodiments. [Figure 6] 1 illustrates a 3D rendition of a portion of an AR or VR device user's face, according to some embodiments. [Figure 7] FIG. 1 is a block diagram of a model architecture used for a 3D rendition of a portion of a VR / AR headset user's face, according to some embodiments. [Figure 8A] 1 illustrates elements and steps in a method for training a model to provide a view of a portion of a user's face on an autostereoscopic display of a virtual reality headset, according to some embodiments. [Figure 8B] 1 illustrates elements and steps in a method for training a model to provide a view of a portion of a user's face on an autostereoscopic display of a virtual reality headset, according to some embodiments. [Figure 8C] 1 illustrates elements and steps in a method for training a model to provide a view of a portion of a user's face on an autostereoscopic display of a virtual reality headset, according to some embodiments. [Figure 8D] 1 illustrates elements and steps in a method for training a model to provide a view of a portion of a user's face on an autostereoscopic display of a virtual reality headset, according to some embodiments. [Figure 9] 1 shows a flowchart of a method for providing an autostereoscopic view of a VR / AR headset user's face, according to some embodiments. [Figure 10] 1 illustrates a flowchart of a method for rendering a three-dimensional (3D) view of a portion of a user's face from multiple two-dimensional (2D) images of the portion of the user's face. [Figure 11] 1 illustrates a flowchart of a method for training a model to render a three-dimensional (3D) view of a portion of a user's face from multiple two-dimensional (2D) images of the portion of the user's face, according to some embodiments. [Figure 12]1 illustrates a computer system configured to implement at least some of the methods for using an AR or VR device, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0026] In the drawings, like elements are labeled likewise according to the description unless otherwise noted.
[0027] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that embodiments of the present disclosure may be practiced without some of these specific details. In other instances, well-known structures and techniques have not been shown in detail so as not to obscure the present disclosure.
[0028] In the field of AR and VR devices and their use, the disconnect that exists between a user and the environment can be annoying to those surrounding the user, if not dangerous to the user or others nearby. In some scenarios, it may be desirable for a user to engage one or more onlookers for conversation or attention. Current AR and VR devices lack the ability to engage onlookers and confirm the focus of the user's attention.
[0029] Typically, display applications that attempt to blend a wide-angle viewing or three-dimensional display with a focal length that provides depth must compromise the spatial resolution of the display. One approach is to reduce the size of the pixels in the display to increase resolution, but current state-of-the-art pixel sizes are approaching the diffraction limit for visible and near-infrared light, limiting the ultimate resolution that can be achieved. For AR and VR devices, this compromise between spatial and angular resolution is less stringent given the limited range of form factors and angular dimensions associated with these devices.
[0030] A desirable feature of AR / VR devices is a small form factor. Therefore, thinner devices are desirable. To achieve this, multi-lenslet array (MLA) light field displays offer a short working distance and a simple design of holographic pancake lenses, providing a thin profile for VR headsets with minimal resolution loss.
[0031] Another desirable feature of AR / VR devices is to provide high resolution, which imposes a limit on the depth of focus, but this limit, common in optical systems used to capture complex scenes, is less severe for the external displays disclosed herein because the depth of field is limited by the relatively constant position between the external display and the user's face.
[0032] The embodiments disclosed herein improve the quality of face-to-face interactions using VR headsets for a wide variety of applications in which one or more people wearing a VR headset interact with one or more people not wearing a VR headset. The embodiments as described herein eliminate friction between the VR user and onlookers or other VR users, bridging the gap between VR and AR, which realizes the benefits of see-through AR coexistence and the enhanced immersion of VR systems. Thus, the embodiments disclosed herein provide an engaging and more natural VR experience.
[0033] More generally, embodiments disclosed herein provide an AR / VR headset that appears to an onlooker like standard see-through glasses, enabling the AR / VR user to better engage with their surroundings, which is extremely useful in scenarios where the AR / VR user interacts with other people or onlookers.
[0034] FIG. 1A illustrates a headset 10A including an autostereoscopic external display 110A according to some embodiments. The headset 10A may be an AR or VR device configured to be worn on a user's head. The headset 10A includes two eyepieces 100A mechanically coupled by straps 15 and having flexible mounts for holding electronic components 20 behind the user's head. A flex connector 5 can electronically couple the eyepieces 100A to the electronic components 20. Each of the eyepieces 100A includes eye imaging systems 115-1 and 115-2 (collectively referred to hereinafter as "eye imaging systems 115") configured to collect images of a portion of the user's face reflected from optical surfaces within a selected field of view (FOV). The eye imaging systems 115 can include dual eye cameras that collect two images of the user's eyes at different FOVs to generate a three-dimensional stereoscopic view of at least a portion of the user's face. The eye imaging systems 115 can provide information regarding pupil position and movement to the electronic components. The eyepiece 100A may also include an external display 110A (e.g., a light field display) adjacent the optical surface and configured to project an autostereoscopic image of the user's face forward from the user.
[0035] In some embodiments, electronic component 20 may include memory circuitry 112 that stores instructions and processor circuitry 122 that executes the instructions to receive images of a portion of the user's face from eye imaging system 115 and provide an autostereoscopic image of the user's face on external display 110A. Additionally, electronic component 20 may also receive images of a portion of the user's face from one or more eye cameras and apply image analysis to assess the user's gaze, vergence, and focus relative to an external scene or aspects of a virtual reality display. In some embodiments, electronic component 20 includes a communications module 118 configured to communicate with a network. Communications module 118 may include radio frequency software and hardware for wirelessly communicating memory 112 and processor 122 with an external network or some other device. Accordingly, communications module 118 may also include a wireless antenna, transceiver, and sensors, as well as digital processing circuitry for signal processing via any one of several wireless protocols, such as Wi-Fi, Bluetooth, or near-field communication (NFC). Additionally, communications module 118 may also communicate with other input tools and accessories (eg, a steering wheel, joystick, mouse, wireless pointer, etc.) that cooperate with headset 10A.
[0036] In some embodiments, eyepiece 100A may include one or more external cameras 125-1 and 125-2 (hereinafter collectively referred to as "external cameras 125") for capturing a front view of a scene for the user. External cameras 125 may be focused or directed (e.g., by processor 122) to aspects of the front view that may be of particular interest to the user based on gaze, convergence, and other characteristics of the user's field of view that may be derived from images of portions of the user's face provided by the dual eye cameras.
[0037] FIG. 1B shows a headset 10B seen by a forward onlooker, according to some embodiments. In some embodiments, the headset 10B can be an AR or VR device in a “snorkel” configuration. Hereinafter, the headsets 10A and 10B are collectively referred to as “headset 10.” In some embodiments, the visor 100B can include a single forward display 110B that provides a view of the user 101 to the onlooker 102. The display 110B includes a portion of the user's 101 face, including both eyes, a portion of the nose, eyebrows, and other facial features. Additionally, an autostereoscopic image 111 of the user's face can include details such as the precise, real-time position of the user's eyes and indicate the user's 101 gaze direction and convergence, or focus of attention. This can indicate to the onlooker 102 whether the user is attending to something being said or to some other environmental disturbance or sensory input that may capture the user's attention.
[0038] In some embodiments, the autostereoscopic image 111 provides a 3D rendering of the user's face. Thus, the spectator 102 has a full-body view of the user's face, and even the user's head, with the viewpoint changing as the spectator 102 changes the angle of view. In some embodiments, the outwardly projected display 110B can include additional image features in the image of a portion of the user's face. For example, in some embodiments, the outwardly projected display can include virtual elements in the image (e.g., a virtual image that the user is actually looking at, or reflections or glare of actual light sources in the environment) superimposed on the image of the user's face.
[0039] FIG. 2 shows a detailed view of an eyepiece 200 for an AR or VR device configured to provide a forward onlooker (see eyepiece 100A and snorkel visor 100B) with a reverse pass-through view of a user's face, according to some embodiments. Eyepiece 200 includes an optical surface 220 configured to provide an image to a user on a first side (left side) of optical surface 220. In some embodiments, the image to the user may be provided by a front camera 225, and optical surface 220 may include a display coupled to front camera 225. In some embodiments, the image in optical surface 220 may be a virtual image provided by a processor executing instructions stored in memory (e.g., memory 112 and processor 122 in the case of a VR device). In some embodiments (e.g., in the case of an AR device), the image to the user may include, at least in part, an image transmitted from the front of eyepiece 200 through a transparent optical component (e.g., a lens, a waveguide, a prism, etc.).
[0040] The eyepiece 200 also includes a first eye camera 215A and a second eye camera 215B (collectively referred to hereinafter as "eye cameras 215") configured to collect first and second images of the user's face (e.g., the user's eyes) at two different FOVs. In some embodiments, the eye camera 215 may be an infrared camera that collects images of the user's face in a reflective mode from the hot mirror assembly 205. The illumination ring 211 can illuminate the portion of the user's face that is to be imaged by the eye camera 215. Accordingly, the optical surface 220 can be configured to be reflective at wavelengths of light operated by the eye camera 215 (e.g., in the infrared region) and transparent at light that provides an image to the user, e.g., in the visible region, including red (R), blue (B), and green (G) pixels. The forward display 210B projects an autostereoscopic image of the user's face to an onlooker (far right of the figure).
[0041] 3A-3D illustrate different aspects and components of a microlens array 300 used as a screen to provide a reverse pass-through view of a user to a forward onlooker in an AR or VR device, according to some embodiments. In some embodiments, the microlens array 300 receives light from the pixel array 320 and provides an image of the user's face to the onlooker. In some embodiments, the image of the user's face is a perspective view of a 3D rendition of the user's face, depending on the onlooker's angle of view.
[0042] 3A is a detailed view of microlens array 300, which includes multiple microlenses 301-1, 301-2, 301-3 (hereinafter collectively referred to as "microlenses 301") arranged in a two-dimensional pattern 302 with pitch 305. In some embodiments, aperture mask 315 is positioned adjacent to the microlens array such that one aperture is aligned with each microlens 301, which can avoid crosstalk of different angles of view from the viewer's perspective.
[0043] For illustrative purposes only, pattern 302 is a hexagonal lattice of microlenses 301 with a pitch 305 of less than 1 millimeter (e.g., 500 μm). Microlens array 300 can include first and second surfaces 310 containing recesses that form microlenses 301, the first and second surfaces 310 being separated by a transparent substrate 307 (e.g., N-BK7 glass, plastic, etc.). In some embodiments, transparent substrate 307 can have a thickness of about 200 μm.
[0044] 3B is a detailed diagram of a light field display 350 for use in a reverse pass-through headset, according to some embodiments. The light field display 350 includes a pixel array 320 adjacent to a microlens array (e.g., microlens array 300), of which only microlens 301 is shown for illustrative purposes. The pixel array 320 includes a plurality of pixels 321 that generate light beams 323 that are directed toward the microlenses 301. In some embodiments, the distance 303 between the pixel array 320 and the microlenses 301 may be approximately equal to the focal length of the microlenses 301; thus, the outgoing light beams 325 may be collimated in different directions depending on the specific location of the outgoing pixels 321. Thus, different pixels 321 in the pixel array 320 can provide different angles of view of a 3D representation of a user's face depending on the position of the onlooker.
[0045] FIG. 3C is a top view of microlens array 300, showing a honeycomb pattern.
[0046] 3D shows microlens array 300 with aperture mask 315 positioned adjacent to it such that the openings in aperture mask 315 are centered on microlens array 300. In some embodiments, aperture mask 315 may comprise chrome with approximately 400 μm openings over a 500 μm hexapack pitch (as shown). Aperture mask 315 may be aligned with first or second surface 310 on either or both sides of microlens array 300.
[0047] FIG. 4 illustrates a ray tracing diagram of a light field display 450 for providing a reverse pass-through image of an AR / VR device user's face to an onlooker, according to some embodiments. The light field display 450 includes a microlens array 400 used to provide a wide-angle, high-resolution view of an AR or VR device user's face to a forward onlooker, according to some embodiments. The microlens array 400 includes a plurality of microlenses 401 arranged in a two-dimensional pattern as disclosed herein. The pixel array 420 may include a plurality of pixels 421 that provide light rays 423 that transmit through the microlens array 400 to generate a 3D rendition of at least a portion of the face of the AR or VR device user. The microlens array 400 may include an aperture mask 415. The aperture mask 415 provides occlusion elements near the edges of each microlens of the microlens array 400. The occlusion elements reduce the amount of light rays 425B and 425C relative to light rays 425A that form a front view of the user's face for the onlooker. This reduces crosstalk and ghosting effects for a viewer positioned in front of the screen and looking at a 3D rendition of the user's face (facing downwards according to Figure 4).
[0048] 5A-5C illustrate different aspects of resolution characteristics 500A, 500B, and 500C (collectively "resolution characteristics 500") in a microlens array for providing a wide-angle, high-resolution view of the face of a user of an AR or VR device, according to some embodiments. The horizontal axis 521 (X-axis) of resolution characteristic 500 represents the image distance (mm) between the user's face (e.g., the user's eyes) and the microlens array. The vertical coordinate 522 (Y-axis) of resolution characteristic 500 represents the resolution of the optical system, including the optical display and screen, given in frequency values, e.g., characteristic cycles on the display per millimeter (cycles / mm), as seen by an onlooker positioned approximately one meter away from the user wearing the AR or VR device.
[0049] FIG. 5A shows resolution characteristics 500A, including a cutoff value, which is the highest frequency a viewer can distinguish from the display. Curves 501-1A and 501-2A (collectively referred to as "curve 501A") are associated with two different headset models (referred to as Model 1 and Model 2, respectively). The specific resolution depends on the image distance and other parameters related to the screen, such as the pitch of the microlens array (e.g., pitch 305). Generally, the greater the distance between the user's eyes and the screen, the monotonically decreasing resolution cutoff (to the right along horizontal axis 521). This is illustrated by the difference between cutoff value 510-2A of curve 501-2A (approximately 0.1 cycles / mm) and cutoff value 510-1A of curve 501-1A (approximately 0.25 cycles / mm). In fact, the image distance for the headset model of curve 501-2A is longer (closer to 10 cm between the user's face and the display) than the headset model of curve 501-1A (approximately 5 cm between the user's eyes and the display). Also, the resolution cutoff is lower for the wide-pitch microlens array (model 2, 500 μm pitch) compared to the narrow-pitch microlens array (model 1, 200 μm pitch).
[0050] FIG. 5B shows a resolution characteristic 500B including curve 501B of a light field display model (Model 3) that provides a spatial frequency of about 0.3 cycles / mm at an image distance of about 5 cm at point 510B.
[0051] FIG. 5C illustrates a resolution characteristic 500C including curves 501-1C, 501-2C, 501-3C, and 501-4C (collectively referred to as "curve 501C"). The horizontal axis 521C (X-axis) of resolution characteristic 500C indicates headset depth (e.g., analogous to the distance between a user's eyes / face and the light field display), and the vertical axis 522C (Y-axis) indicates the pixel pitch (in microns, μm) of the pixel array within the light field display. Each of curves 501C indicates the cycles / mm cutoff resolution for each light field display model. Point 510B is shown in comparison to the better resolution achieved at point 510C for a light field display model (Model 4) with dense pixel packing (pitch less than 10 μm) and a closer headset depth of approximately 25 mm (e.g., approximately 1 inch or less).
[0052] 5D shows images 510-1D and 510-2D of a user wearing a headset by a spectator for each of the light field display models. Image 510-1D is obtained with Model 3, and image 510-2D is obtained with Model 4 (see points 510B and 510C, respectively). The resolution performance of Model 4 is certainly better than that of Model 3, demonstrating the wide range of possibilities for accommodating desired resolutions after considering other tradeoffs in model design, consistent with this disclosure.
[0053] FIG. 6 illustrates 3D renditions 621A and 621B (collectively referred to hereinafter as “3D renditions 621”) of a portion of a face of a user of an AR or VR device, according to some embodiments. In some embodiments, 3D rendition 621 is provided by a model 650 operating on multiple 2D images 611 of at least a portion of the user's face (e.g., eyes) and provided by an eye imaging system of the AR or VR device (see eye imaging system 115 and eye camera 215). Model 650 may include linear and / or nonlinear algorithms, such as neural networks (NNs), convolutional neural networks (CNNs), machine learning (ML) models, and artificial intelligence (AI) models. Model 650 includes instructions stored in a memory circuit and executed by a processor circuit. The memory circuit and processor circuit may be stored on the back of the AR or VR device (e.g., memory 112 and processor 122 in electronics 20). Thus, multiple 2D images 611 are received from the eye imaging system to create, update, and improve model 650. The multiple 2D images include at least two different FOVs, e.g., from each of two different stereoscopic cameras in the eye imaging system, and model 650 can determine which images came from which cameras to form 3D rendition 621. Model 650 then uses the 2D image input and detailed knowledge of the difference between the FOVs of the two eye cameras (e.g., camera direction vectors) to provide 3D rendition 621 of at least a portion of the face of a user of the AR or VR device.
[0054] 7 shows a block diagram of a model architecture 700 used for 3D rendition of a VR / AR headset user's facial region, according to some embodiments. The model architecture 700 is a pixel-aligned volumetric avatar (PVA) model. The PVA model 700 is trained from a multi-perspective image collection that produces multiple 2D input images 701-1, 701-2, and 701-n (collectively referred to hereafter as "input images 701"). Each of the input images 701 is associated with a camera view vector v that indicates the viewing direction of the user's face for that particular image. i (e.g., v1, v2, and v n ) is related to the vector v i Each of these is a camera-specific parameter K i and rotation R i (For example, {K i , [R|t] i}) are known viewpoints 711 associated with the camera. i can include brightness, color mapping, sensor efficiency, and other camera-dependent parameters. i indicates the orientation (and distance) of the subject's head relative to the camera. Different camera sensors will respond slightly differently to the same incident radiance, despite the fact that they are the same camera model. If nothing is done to address this, the intensity differences will eventually be baked into the scene representation N, causing the image to appear unnaturally bright or dark from certain viewpoints. To address this, bias and gain values for each camera are learned. This allows the system to have a "simpler" way of accounting for variability in the data.
[0055] The value of "n" is purely exemplary, and one skilled in the art will understand that any number of n input images 701 can be used. The PVA model 700 generates a volumetric rendition 721 of the headset user. The volumetric rendition 721 is a 3D model (e.g., an "avatar") that can be used to generate a 2D image of the object from the target's perspective. This 2D image changes as the target's perspective changes (e.g., as an onlooker moves around the headset user).
[0056] The PVA model 700 includes a convolutional encoder-decoder 710A, a ray marching stage 710B, and a radiance field stage 710C (collectively referred to hereinafter as "PVA stage 710"). The PVA model 700 is trained on input images 701 selected from a multi-identity training corpus using a steepest descent method. Thus, the PVA model 700 includes a loss function defined between predicted images from multiple objects and corresponding ground truth. This enables the PVA model 700 to render accurate volumetric renditions 721 regardless of the object.
[0057] A convolutional encoder-decoder network 710A takes in an input image 701 and produces pixel-aligned feature maps 703-1, 703-2, and 703-n (collectively referred to as "feature maps 703"). A ray marching stage 710B tracks each pixel along a ray of target view j defined by {Kj, [R|t]j} and accumulates the color, c, and optical density ("opacity") produced by a radiance field stage 710C at each point. A radiance field stage 710C(N) converts the 3D positions and pixel-aligned features to color and opacity and renders a radiance field 715(c, σ).
[0058] The input image 701 is i721-n (hereafter collectively referred to as "learnable weights 721"). The ray marching stage 710B performs world-to-camera projection 723, bilinear interpolation 725, position encoding 727, and feature aggregation 729.
[0059] In some embodiments, the conditioning view v i ∈R h×w×3 Then, the feature map 703 may be defined as a function as follows: TIFF0007825624000001.tif11170
[0060] where φ(X):R 3 →R 6×l is a set of points 730 (X∈R) with 2×l different basis functions. 3 ) is an encoding of the location of the point 730 (X) along a ray directed from the 2D image of the object to a particular viewpoint 731, r0. (i) ∈R h×w×d ) is the camera position vector v i where d is the number of feature channels, h and w are the height and width of the image, and f X ∈R d are the aggregated image features associated with point X. Each feature map f (i) , the ray marching stage 710B computes f by projecting a 3D point X along a ray using the camera's intrinsic (K) and extrinsic (R, t) parameters for that particular viewpoint. X ∈R d get. TIFF0007825624000002.tif15170
[0061] where Π is the perspective projection function to the camera pixel coordinates, and F(f, x) is the bilinear interpolation 725 of f at pixel location x. The ray marching stage 710B generates pixel-aligned features f from multiple images for the radiance field stage 710C. (i) X Combine.
[0062] Camera intrinsic function K j , rotation and translation R j , t j Given a training image v with j For camera and center 731(r0)∈R 3 pixel p∈R for a given viewpoint in the focal plane of 2 The predicted color of is determined by the "camera-to-world projection matrix" P -1 =[R i |t i ] -1 K -1 i is obtained by marching a ray into the scene using the , and the direction of the ray is given by TIFF0007825624000003.tif8170
[0063] The ray marching stage 710B calculates the ray vectors t∈[t near , t far ] along the ray 735 defined by r(t)=r0+td, accumulate radiance and opacity values as follows: TIFF0007825624000004.tif7170
[0064] where: TIFF0007825624000005.tif8170
[0065] In some embodiments, the ray marching stage 710B includes n s A set of points t~[t near , t far ] uniformly. By setting X = r(t), we can use quadrature rules to approximate integrals 6 and 7. Function Iα (p) can be defined as follows: TIFF0007825624000006.tif7170
[0066] where α i =1-exp(-δ i σ i ) and δ i is the distance between the i+1th and ith sample points along the ray 735.
[0067] Known camera viewpoints v i In a multi-view setting with a fixed number of conditioning views, the ray marching stage 710B aggregates features by simple concatenation. i} n i=1 and n conditioning images {v i} n i=1 For each point X, we define the feature {f (i) x} n i=1 Using this, the final features are generated by ray marching stage 710B as follows: TIFF0007825624000007.tif7170
[0068] where: TIFF0007825624000008.tif6170 represents connectivity along the depth dimension, which is the feature information from the viewpoint. n i=1 This saves the data and helps the PVA Model 700 determine the best combination and employ conditioning information.
[0069] In some embodiments, the PVA model 700 is independent of the number of viewpoints and conditioning views. A simple concatenation as described above is insufficient in this case because the number of conditioning views is not known in advance, which may result in different feature dimensions (d) during inference time. To summarize the characteristics of the multi-view setting, some embodiments define a permutation-invariant function G:R such that for any permutation ψ, n×d →R d Includes. G(f (1) ...,f (n) )=G([f ψ(1) ,f ψ(2) ...,f ψ(n) ])
[0070] A simple permutation-invariant function for feature aggregation is the average of the sampled feature maps 703. This aggregation procedure may be desirable in cases where depth information is available during training. However, in cases where there is depth ambiguity (e.g., for points projected onto the feature maps 703 before sampling), the above aggregation may result in artifacts. To avoid this, some embodiments consider camera information to include effective conditioning in the radiance field stage 710C. Thus, some embodiments consider the feature vector f (i) X , and camera information (ci) are taken into account to generate a camera summary feature vector f' (i) X A conditioning function network N that generates cf :R d+7 →R d’ These modified vectors are then averaged over multiple or all conditioning views as follows: TIFF0007825624000009.tif20170
[0071] The advantage of this approach is that the features summarized by the camera can take into account possible occlusions before feature averaging is performed. The camera information is encoded as a 4D rotation quaternion and a 3D camera position.
[0072] Some embodiments also use a background estimation network N to avoid learning part of the background in the scene representation. bg It is also possible to include a background estimation network N bg is N bg :R nc :→R h×w×3 and learns a fixed background for each camera. In some embodiments, the radiance field stage 710C learns a fixed background for each camera. bg can be used to predict the final image pixels as follows: I p =I rgb +(1-Iα)·I bg (11)
[0073] Camera C i against TIFF0007825624000010.tif6170, where TIFF0007825624000011.tif6170 is an initial estimate of the background extracted using Inpaint. α is defined in equation (8). These filled backgrounds are often noisy, resulting in a "halo" effect around the person's head. To avoid this, N bg The model learns residuals against the inpainted background, which has the advantage of not requiring a large network to account for the background.
[0074] Target ground truth image v j For, the PVA model 700 uses a simple photometric reconstruction loss to train both the radiance field stage 710C and the feature extraction network. TIFF0007825624000012.tif8170
[0075] 8A-8D illustrate elements and steps in a method for training a model to provide a view of a portion of a user's face on an autostereoscopic display in a virtual reality headset, according to some embodiments. Eyepieces 800 are trained with multiple training images 811 from multiple users. A 3D model 821 is created for each user, including texture and depth maps (collectively referred to hereinafter as "texture-depth maps 833") to recover details of image features 833-1B, 833-2B, and 833C. Once the 3D models 821 are generated, an autostereoscopic image of the three-dimensional reconstruction of the user's face is provided to a pixel array of a light field display. The light field display is separated into multiple segments of active pixels, each segment providing a portion of the field of view of the 3D model 821 at a selected angle of view for an onlooker.
[0076] FIG. 8A shows a setup 850 for collecting multiple training images 811 on an eyepiece 800, according to some embodiments. The training images 811 may be provided by a display and projected onto a screen 812 located in the same position as the hot mirror would be when the eyepiece is assembled in the headset. One or more infrared cameras 815 collect the training images 811 in a reflection mode, and one or more RGB cameras 825 collect the training images in a transmission mode. The setup 850 has an image vector 801-1, an IR camera vector 801-2, and an RGB camera vector 801-3 (collectively referred to hereafter as "positioning vector 801") fixed for all training images 811. The positioning vector 801 is used by an algorithmic model to accurately estimate the size, distance, and angle of view associated with a 3D model 821.
[0077] 8B shows a texture image 833-1B and a depth image 833-2B, according to some embodiments. Texture image 833-1B may be obtained from capturing training images using RGB camera 825, and depth image 833-2B may be obtained from training images using IR camera 815.
[0078] Figure 8C shows a depth image 833C collected by the IR camera 815, according to some embodiments. Figure 8D shows a 3D model 821 formed in relation to the eyepiece 800, according to some embodiments.
[0079] 9 shows a flowchart of a method 900 for providing an autostereoscopic view of a VR / AR headset user's face, according to some embodiments. The steps of method 900 may be performed at least in part by a processor executing instructions stored in memory, where the processor and memory are part of the headset electronics disclosed herein (e.g., memory 112, processor 122, electronics 20, and headset 10). In still other embodiments, at least one or more of the steps of a method consistent with method 900 may be performed by a processor executing instructions stored in memory, where at least one of the processor and memory is remotely located in a cloud server, and the headset device is communicatively coupled to the cloud server via a network-coupled communications module (see communications module 118). In some embodiments, method 900 can be implemented using a model including a neural network architecture in a machine learning or artificial intelligence algorithm, as disclosed herein (e.g., model 650, model architecture 700). In some embodiments, methods consistent with the present disclosure may include at least one or more steps from method 900 performed in a different order, simultaneously, quasi-simultaneously, or overlapping in time.
[0080] Step 902 includes receiving a plurality of images from one or more headset cameras having at least two or more fields of view of a subject, the headset user.
[0081] Step 904 includes extracting a plurality of image features from the image using a set of learnable weights. In some embodiments, step 904 includes matching image features along scanlines to construct a cost volume at a first resolution setting and provide a coarse disparity estimate. In some embodiments, step 904 includes recovering one or more image features, including small details or thin structures, at a second resolution setting that is higher than the first resolution setting. In some embodiments, step 904 includes generating a texture map of portions of the user's face and a depth map of portions of the user's face based on the image features, wherein the texture map includes color details of the image features and the depth map includes depth locations of the image features. In some embodiments, step 904 includes extracting unique characteristics of the headset camera used to collect each image.
[0082] Step 906 involves forming a three-dimensional model of the object using the learnable weights.
[0083] Step 908 includes mapping the three-dimensional model of the object to an autostereoscopic display format that associates an image projection of the object with a selected observation point for the spectator. In some embodiments, step 908 includes providing a portion of a field of view of a user's face at a selected viewpoint for the spectator in one segment of the light field display. In some embodiments, step 908 further includes tracking one or more spectators to identify an angle of view and modifying the light field display to optimize the field of view for each of the one or more spectators. In some embodiments, step 908 includes interpolating a feature map associated with a first observation point with a feature map associated with a second observation point. In some embodiments, step 908 includes aggregating image features of multiple pixels along a direction of the selected observation point. In some embodiments, step 908 includes concatenating multiple feature maps generated by each headset camera in a permutation-invariant combination, each headset camera having unique characteristics.
[0084] Step 910 includes providing an image projection of the object on the display when the spectator is located at the selected observation point. In some embodiments, step 910 includes providing a second image projection on the device display as the spectator moves from the first observation point to the second observation point.
[0085] 10 shows a flowchart of a method 1000 for rendering a three-dimensional (3D) view of a portion of a user's face from multiple two-dimensional (2D) images of the portion of the user's face. The steps of method 1000 may be performed at least in part by a processor executing instructions stored in memory, where the processor and memory are part of the electronic components of a headset disclosed herein (e.g., memory 112, processor 122, electronic components 20, and headset 10). In still other embodiments, at least one or more of the steps in a method consistent with method 1000 may be performed by a processor executing instructions stored in memory, where at least one of the processor and memory is remotely located in a cloud server, and the headset device is communicatively coupled to the cloud server via a communication module (see communication module 118) coupled to a network. In some embodiments, method 1000 may be performed using a model (e.g., model 650, model architecture 700) that includes a neural network architecture in a machine learning or artificial intelligence algorithm, as disclosed herein. In some embodiments, methods consistent with the present disclosure may include at least one or more steps from method 1000 performed in a different order, simultaneously, quasi-simultaneously, or overlapping in time.
[0086] Step 1002 involves collecting a plurality of ground truth images of a plurality of users' faces.
[0087] Step 1004 includes correcting the ground truth image using a stored, calibrated pair of stereoscopic images. In some embodiments, step 1004 includes extracting a plurality of image features from the two-dimensional image using a set of learnable weights. In some embodiments, step 1004 includes extracting intrinsic characteristics of a camera used to collect the two-dimensional image.
[0088] Step 1006 includes mapping the three-dimensional model of the object to an autostereoscopic display format that associates an image projection of the object with a selected observation point for the spectator. In some embodiments, step 1006 includes projecting image features along a direction between the three-dimensional model of the object and the selected observation point for the spectator. In some embodiments, step 1006 includes interpolating a feature map associated with a first direction with a feature map associated with a second direction. In some embodiments, step 1006 includes aggregating image features for a plurality of pixels along a direction between the three-dimensional model of the object and the selected observation point. In some embodiments, step 1006 includes concatenating multiple feature maps generated by each of a plurality of cameras in a permutation-invariant combination, each of the multiple cameras having unique characteristics.
[0089] Step 1008 includes determining a loss value based on a difference between the ground truth image and the image projection of the object. In some embodiments, step 1008 includes providing an autostereoscopic image of a three-dimensional model of the object to an onlooker. In some embodiments, step 1008 includes evaluating a loss function based on a difference between the autostereoscopic image of the three-dimensional model of the object and the ground truth image of the object, and updating at least one of the sets of learnable weights based on the loss function.
[0090] Step 1010 involves updating the 3D model of the object based on the loss value.
[0091] 11 shows a flowchart of a method 1100 for training a model that renders a three-dimensional (3D) view of a portion of a user's face from multiple two-dimensional (2D) images of the portion of the user's face, according to some embodiments. The steps of method 1100 may be performed at least in part by a processor executing instructions stored in memory, where the processor and memory are part of the electronic components of a headset disclosed herein (e.g., memory 112, processor 122, electronic components 20, and headset 10). In still other embodiments, at least one or more of the steps in a method consistent with method 1100 may be performed by a processor executing instructions stored in memory, where at least one of the processor and memory is remotely located in a cloud server, and the headset device is communicatively coupled to the cloud server via a communication module (see communication module 118) coupled to a network. In some embodiments, method 1100 can be implemented using a model (e.g., model 650, model architecture 700) that includes a neural network architecture in a machine learning or artificial intelligence algorithm, as disclosed herein. In some embodiments, methods consistent with the present disclosure may include at least one or more steps from method 1100 performed in a different order, simultaneously, quasi-simultaneously, or overlapping in time.
[0092] Step 1102 includes collecting a plurality of ground truth images of a plurality of users' faces.
[0093] Step 1104 involves correcting the ground truth image using the stored calibrated stereoscopic image pair.
[0094] Step 1106 includes generating multiple synthetic views of the object using the 3D face model, where the synthetic views of the object include an interpolation of multiple feature maps projected along different directions corresponding to the multiple views of the object. In some embodiments, step 1106 includes projecting image features from each of the ground truth images along selected viewing directions and concatenating the multiple feature maps generated by each of the ground truth images in a permutation-invariant combination, where each of the ground truth images has unique characteristics.
[0095] Step 1108 includes training a three-dimensional face model based on differences between the ground truth image and the synthetic view of the object. In some embodiments, step 1108 includes updating at least one in the set of learnable weights for each of the plurality of features in the feature map based on a value of a loss function indicative of the difference between the ground truth image and the synthetic view of the object. In some embodiments, step 1108 includes training a background value for each of the plurality of pixels in the ground truth image based on pixel background values projected from the plurality of ground truth images.
[0096] Hardware Overview 12 is a block diagram illustrating an exemplary computer system 1200 on which headset 10 and methods 900, 1000, and 1100 can be implemented. In certain aspects, computer system 1200 may be implemented using hardware or a combination of software and hardware, either on a dedicated server, integrated into another entity, or distributed across multiple entities. Computer system 1200 may include a desktop computer, a laptop computer, a tablet, a phablet, a smartphone, a feature phone, a server computer, etc. The server computer may be located remotely in a data center or may be stored locally.
[0097] Computer system 1200 includes a bus 1208 or other communication mechanism for communicating information, and a processor 1202 (e.g., processor 122) coupled to bus 1208 for processing information. By way of example, computer system 1200 may be implemented with one or more processors 1202. Processor 1202 may be a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gated logic, a discrete hardware component, or any other suitable entity capable of performing calculations or other manipulations of information.
[0098] Computer system 1200 may include, in addition to hardware, one or more combinations of code that create an execution environment for the computer programs, such as processor firmware, protocol stacks, database management systems, code constituting an operating system, or code stored in included memory 1204 (e.g., memory 112), such as random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable PROM (EPROM), registers, hard disk, removable disk, CD-ROM, DVD, or any other suitable storage device coupled to bus 1208 for storing information and instructions executed by processor 1202. Processor 1202 and memory 1204 may be supplemented by, or incorporated in, special purpose logic circuitry.
[0099] The instructions may be stored in memory 1204 and implemented in one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of computer system 1200, and may be implemented according to any method known to those skilled in the art, including, but not limited to, computer languages such as data-oriented languages (e.g., SQL, dBase), system languages (e.g., C, Objective-C, C++, Assembly), architecture languages (e.g., Java, .NET), and application languages (e.g., PHP, Ruby, Perl, Python). The instructions may also be implemented in a computer language such as an array language, an aspect-oriented language, an assembly language, an authoring language, a command line interface language, a compiled language, a parallel language, a curly bracket language, a dataflow language, a data structuring language, a declarative language, an esoteric language, an extensible language, a fourth generation language, a functional language, an interactive mode language, an interpretive language, an iterative language, a list-based language, a little language, a logic-based language, a machine language, a macro language, a metaprogramming language, a multi-paradigm language, numerical analysis, a non-English-based language, an object-oriented class-based language, an object-oriented prototype-based language, an offside rules language, a procedural language, a reflective language, a rule-based language, a scripting language, a stack-based language, a synchronous language, a syntax processing language, a visual language, a worst-case language, and an XML-based language. The memory 1204 may also be used for storing temporary variables or other intermediate information during execution of instructions by the processor 1202.
[0100] The computer programs described herein do not necessarily correspond to files in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be executed on one computer or deployed to run on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network. The processes and logic flows described herein can be implemented by one or more programmable processors executing one or more computer programs, operating on input data and generating output to perform functions.
[0101] Computer system 1200 further includes a data storage device 1206, such as a magnetic or optical disk, coupled with bus 1208 for storing information and instructions. Computer system 1200 may be coupled to various devices via input / output module 1210. Input / output module 1210 may be any input / output module. Exemplary input / output module 1210 includes a data port, such as a USB port. Input / output module 1210 is configured to connect to a communications module 1212. Exemplary communications module 1212 includes a networking interface card, such as an Ethernet card and a modem. In certain aspects, input / output module 1210 is configured to connect to multiple devices, such as input devices 1214 and / or output devices 1216. Exemplary input devices 1214 include a keyboard and a pointing device, such as a mouse or trackball, that allows a consumer to provide input to computer system 1200. Other types of input devices 1214 can also be used to provide interaction with the consumer, such as tactile input devices, visual input devices, audio input devices, or brain-computer interface devices. For example, feedback provided to the consumer can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the consumer can be received in any form, including acoustic input, audio input, tactile input, or brainwave input. Exemplary output devices 1216 include a display device, such as an LCD (liquid crystal display) monitor, for displaying information to the consumer.
[0102] According to one aspect of the present disclosure, headset 10 may be implemented at least in part using computer system 1200 in response to processor 1202 executing one or more sequences of one or more instructions contained in memory 1204. Such instructions may be read into memory 1204 from another machine-readable medium, such as data storage device 1206. Execution of the sequences of instructions contained in main memory 1204 causes processor 1202 to perform the process steps described herein. One or more processors in a multi-processing configuration may also be employed to execute the sequences of instructions contained in memory 1204. In alternative aspects, hardwired circuitry may be used in place of or in combination with software instructions to implement various aspects of the present disclosure. Thus, aspects of the present disclosure are not limited to any specific combination of hardware circuitry and software.
[0103] Various aspects of the subject matter described herein may be implemented in a computing system that includes back-end components such as a data server, or that includes middleware components such as an application server, or that includes front-end components, e.g., a client computer having a graphical consumer interface or web browser through which a consumer can interact with an implementation of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. The communications network may include, for example, any one or more of a LAN, a WAN, the Internet, etc. Furthermore, the communications network may include, for example, but is not limited to, any one or more of the following network topologies, including, for example, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or hierarchical network, etc. The communications module may be, for example, a modem or an Ethernet card.
[0104] Computer system 1200 may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers having a client-server relationship to each other. Computer system 1200 may be, for example, without limitation, a desktop computer, a laptop computer, or a tablet computer. Computer system 1200 may also be incorporated into another device, for example, without limitation, a mobile phone, a PDA, a mobile audio player, a global positioning system (GPS) receiver, a video game console, and / or a television set-top box.
[0105] The terms "machine-readable storage medium" or "computer-readable medium," as used herein, refer to any one or more media that participate in providing instructions to processor 1202 for execution. Such media may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as data storage device 1206. Volatile media include dynamic memory, such as memory 1204. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that form bus 1208. Common forms of machine-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, any other optical media, punch cards, paper tape, any other physical media with a pattern of holes, RAM, PROM, EPROM, FLASH EPROM, any other memory chip or cartridge, or any other medium from which a computer can read. The machine-readable storage medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that affects a machine-readable propagated signal, or a combination of one or more thereof.
[0106] To illustrate the interchangeability of hardware and software, various illustrative blocks, modules, components, methods, operations, instructions, algorithms, etc. have been described generally in terms of their functionality. Whether such functionality is implemented as hardware, software, or a combination of hardware and software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application.
[0107] As used herein, the terms "and" or "or" separate any of the items, and the phrase "at least one" preceding a series of items modifies the entire list, not each member (e.g., each item) of the list. The phrase "at least one" does not require the selection of at least one item; rather, the phrase allows for the meaning of including at least one of any one of the items, and / or at least one of any combination of the items, and / or at least one of each of the items. By way of example, the phrases "at least one of A, B, and C" or "at least one of A, B, or C" refer to A only, B only, or C only, any combination of A, B, and C, and / or at least one of each of A, B, and C, respectively.
[0108] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. Phrases such as "one aspect," "that aspect," "another aspect," "some aspects," "one or more aspects," "one implementation," "that implementation," "another implementation," "some implementations," "one or more implementations," "one embodiment," "that embodiment," "another embodiment," "some embodiments," "one or more embodiments," "one configuration," "one configuration," "another configuration," "some configurations," "one or more configurations," the subject technology, the disclosure, the present disclosure, and other variations thereof, are used for convenience and do not imply that disclosure associated with such phrases is essential to the subject technology or that such disclosure applies to all configurations of the subject technology. Disclosures associated with such phrases may apply to all configurations or to one or more configurations. Disclosures associated with such phrases may provide one or more examples. Phrases such as "one aspect" or "some aspects" can refer to one or more aspects, and vice versa, as apply to other such phrases.
[0109] Reference to an element in the singular does not mean "one and only one" unless specifically stated otherwise, but rather means "one or more." The term "some" refers to one or more. Underlined and / or italicized headings and subheadings are used for convenience only and do not limit the subject technology, and are not to be construed in connection with the interpretation of the description of the subject technology. Relative terms such as "first," "second," etc. may be used to distinguish one entity or act from another, but do not necessarily require or imply such an actual relationship or order between such entities or acts. All structural and functional equivalents to the elements of the various configurations described throughout this disclosure that are or later become known to those skilled in the art are encompassed by the subject technology. Furthermore, nothing disclosed herein is intended to be publicly available, regardless of whether such disclosure is expressly stated in the description above. No element of a claim shall be construed under the provisions of 35 U.S.C. § 112, sixth paragraph, unless the element is expressly recited using the phrase "means for," or, in the case of a method claim, unless the element is recited using the phrase "step for."
[0110] While this specification contains many details, these should not be construed as limitations on the scope of what may be described, but rather as descriptions of particular implementations of the subject matter. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features are described above as acting in particular combinations, and may even be initially described as such, in some cases one or more features from a described combination can be excluded from that combination, and a described combination may be directed to a subcombination or a variation of the subcombination.
[0111] Although the subject matter herein has been described in particular embodiments, other embodiments may be implemented and are within the scope of the following claims. For example, while acts are depicted in the figures in a particular order, this should not be understood as requiring that such acts be performed in the particular order or sequential order depicted, or that all of the acts depicted be performed, to achieve desired results. Actions recited in the claims may be performed in a different order and still achieve desirable results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order or sequential order depicted to achieve desired results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.
[0112] The title, background, brief description of the drawings, abstract, and drawings are hereby incorporated into this disclosure and are provided as illustrative examples of the disclosure, not as a limiting description. They are submitted with the understanding that they will not be used to limit the scope or meaning of the claims. Furthermore, in the detailed description, it will be appreciated that the description provides illustrative examples, and that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the described subject matter requires more features than are expressly recited in each claim. Rather, as the claims reflect, inventive subject matter lies in less than all features of a single disclosed structure or operation. Accordingly, the claims are incorporated into the detailed description, with each claim standing on its own as separately described subject matter.
[0113] The claims are not intended to be limited to the embodiments described herein, but are to be accorded the full scope consistent with the language of the claims and encompass all legal equivalents. Nonetheless, none of the claims should be construed to encompass subject matter that does not comply with applicable patent law requirements.
Claims
1. A device, a near-eye display configured to provide an image to a subject; an eye imaging system configured to acquire images of the subject; a light field display configured to provide an autostereoscopic image of a three-dimensional model of the object to a viewer, the autostereoscopic image including perspective-corrected views of the object from multiple viewpoints within a field of view of the light field display; and A device comprising:
2. 2. The device of claim 1, wherein the light field display includes a pixel array and a multi-lenslet array, the pixel array configured to provide a segmented view of the object to the multi-lenslet array, the segmented view including multiple portions of a field of view of the light field display at a selected viewpoint.
3. 3. The device of claim 1 or 2, wherein the eye imaging system comprises two cameras for collecting a binocular view of the subject.
4. and / or The device of any one of claims 1 to 3, wherein the near-eye display provides the subject with a three-dimensional representation of an environment including the spectator.
5. 5. The device of claim 1, wherein the eye imaging system includes an infrared camera that receives the image from the subject in a reflection mode from a dichroic mirror adjacent to the light field display.
6. A device described in any one of claims 1 to 5, wherein the light field display comprises a microlens array having a plurality of microlenses arranged in a two-dimensional pattern having a preselected pitch to avoid crosstalk between the perspective-corrected views for two viewpoints for the spectator.
7. 7. The device of claim 1, wherein the light field display further comprises an aperture mask adjacent to the microlens array, the aperture mask including a plurality of apertures, each aperture aligned with the center of a respective microlens in the microlens array.
8. the light field display includes a pixel array divided into a plurality of active segments, each active segment in the pixel array having a dimension corresponding to a diameter of a refractive element in a multi-lenslet array; and / or The device of any one of claims 1 to 7, further comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the light field display to divide a pixel array into a plurality of active segments, each active segment configured to provide a portion of the field of view of the light field display at a selected viewpoint for the spectator.
9. 1. A computer-implemented method comprising: receiving a plurality of images from one or more headset cameras, the plurality of images having at least two or more fields of view of a subject, the headset user; extracting a plurality of image features from the image using a set of learnable weights; forming a three-dimensional model of the object using the set of learnable weights; and mapping the three-dimensional model of the object into an autostereoscopic display format that associates an image projection of the object with a selected observation point for an onlooker; providing the image projection of the object on a device display when the spectator is located at the selected observation point; 11. A computer-implemented method comprising:
10. The computer-implemented method of claim 9 , wherein extracting image features comprises extracting unique characteristics of a headset camera used to capture each of the images.
11. Mapping the three-dimensional model of the object into an autostereoscopic display format comprises interpolating a feature map associated with a first observation point with a feature map associated with a second observation point; or Mapping the three-dimensional model of the object to an autostereoscopic display format includes aggregating the image features for a plurality of pixels along a direction of the selected observation point; or 11. The computer-implemented method of claim 9 or claim 10, wherein mapping the three-dimensional model of the object to an autostereoscopic display format includes concatenating multiple feature maps generated by each of the headset cameras in a permutation-invariant combination, each of the headset cameras having unique characteristics.
12. 12. The computer-implemented method of claim 9, wherein providing the image projection of the object comprises providing a second image projection on the device display as the spectator moves from a first observation point to a second observation point.
13. 1. A computer-implemented method for training a model to provide views of an object on an autostereoscopic display in a virtual reality headset, comprising generating a plurality of synthetic views, collecting a plurality of ground truth images of a plurality of users' faces; correcting the ground truth image using stored calibrated stereoscopic image pairs; mapping the three-dimensional model of the object into an autostereoscopic display format that associates an image projection of the object with a selected observation point for an onlooker; determining a loss value based on a difference between the ground truth image and the image projection of the object; updating the three-dimensional model of the object based on the loss value; 11. A computer-implemented method comprising:
14. Generating a plurality of synthetic views includes projecting image features from each of the ground truth images along a selected observation direction; and concatenating a plurality of feature maps generated by each of the ground truth images in a permutation-invariant combination, each of the ground truth images having unique characteristics; and / or training the three-dimensional model of the object includes updating at least one of a set of learnable weights for each of a plurality of features based on a value of a loss function indicative of the difference between the ground truth image and the image projection of the object; or 14. The computer-implemented method of claim 13, wherein training the three-dimensional model of the object comprises training a background value for each of a plurality of pixels in the ground truth image based on pixel background values projected from the plurality of ground truth images.
Citation Information
Patent Citations
Light field projector based on movable LED array and microlens array for use in head-mounted displays
JP2015521298A
Display device and display method
JP2016109760A
Information processing program, information processing method and program
JP2017069687A
Head mounted display device, program, and control method of head mounted display device
JP2018142857A
Head-Mounted Light Field Display Using Integral Imaging and Waveguide Prisms
JP2020514811A