Eye tracking assisted perspective rendering
By using eye tracking to generate a constant depth grid in the MR system, the problem of inaccurate depth representation in perspective rendering is solved, achieving a balance between low computational cost and visual accuracy, and reducing visual artifacts.
Patent Information
- Application Number
- CN202510267507.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2025-03-07
- Publication Date
- 2025-09-23
AI Technical Summary
Existing MR systems have difficulty generating accurate depth representation in perspective rendering, resulting in visual artifacts and computational burden, especially in the user's viewing area, and are unable to balance perceptual accuracy with system power and latency constraints.
By using eye tracking information to determine the three-dimensional position where the user is looking, a depth mesh with constant depth is generated, and the image captured by the camera is reprojected into the user's eye space, balancing computational cost and visual accuracy.
The generated perspective scene is accurate and distortion-free in the user's area of interest, with low computational cost, reduced visual artifacts, and meets the system's power and latency requirements.
Smart Images

Figure CN120689491A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. non-provisional patent application No. 18 / 612,041, filed on March 21, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to systems and methods designed for immersive rendering of mixed-reality (MR) scenes for users. Background Art
[0004] A head-mounted device (HMD) featuring a stereoscopic display can provide an immersive experience in a three-dimensional environment. When a user wears an HMD, their vision of the surrounding physical environment is obscured by the HMD's physical structure and display. Mixed reality (MR) addresses this problem by using the HMD's camera to capture a real-time, low-latency live feed of the surrounding physical environment and displaying this live feed to the user, allowing the user to seamlessly perceive their environment as if they were not wearing an HMD. In addition, users can enhance their environment by overlaying virtual elements on the real world.
[0005] "See-through" refers to the MR feature that allows the user to see their physical environment while wearing an HMD. Information about the user's physical environment is visually "see-through" to the user by having the artificial reality system's headset display information captured by the headset's outward-facing cameras. Simply displaying the captured images will not work. Because the position of the camera does not align with the position of the user's eyes, the images captured by the cameras do not accurately reflect the user's perspective. In addition, because the images have no depth information, simply displaying the images will not provide the user with the appropriate parallax effect if the user moves away from the location where the image was captured. Incorrect parallax combined with user motion may cause motion sickness.
[0006] Perspective images are generated by reprojecting or warping images captured by the cameras of the artificial reality device toward the user's eye position using depth measurements of the scene (depth may be measured using a depth sensor and / or machine learning-based methods). The artificial reality head-mounted view device may have an outward-facing left camera and an outward-facing right camera, which are used to capture images for perspective generation. Based on the depth estimation of the scene, the left image captured by the left camera is reprojected to the viewpoint of the left eye, and the right image captured by the right camera is reprojected to the viewpoint of the right eye. When the reprojected images captured by the cameras are displayed to the user, these images will approximate what the captured scene would look like from the perspective of the user's two eyes.
[0007] Since reprojection relies on depth information of the physical scene, the accuracy of the depth representation (e.g., depth mesh) plays an important role. In practice, it is difficult to generate pixel-accurate depth representation for the entire visible scene in real time. Not only is high-resolution and accurate depth sensing challenging, but it also needs to be robust enough to adapt to different lighting conditions, object movement, head motion, occlusion, and other environmental factors. In addition, generating a depth representation of the scene from the acquired depth information can be computationally expensive. In the context of perspective generation for MR, depth sensing and the generation of depth representation need to be implemented under strict timing constraints, limited power budget, and increased accuracy requirements. Therefore, designing a suitable technique to generate depth representation for perspective rendering is a complex challenge for developers and researchers.
[0008] Some existing systems address the above challenges by approximating the depth of a scene using a continuous, spatially varying depth mesh shaped according to a general outline of the scene depth. This depth mesh is like a blanket thrown over the physical objects in the scene. The benefit of this depth mesh is that it balances the trade-off between acquiring scene depth information and computational complexity. However, a disadvantage is that a continuous depth mesh may have multiple areas of inaccurate depth. For example, if the physical environment includes a foreground object and a background object, the depth mesh may approximate the actual depth of these objects fairly well. However, because the depth mesh is continuous (e.g., like a blanket), the area between these two objects in the depth mesh will be inaccurate. In turn, the inaccuracies in the depth mesh will lead to inaccurate reprojections of the perspective image. The end result is that the perspective image will exhibit visual artifacts in the form of deformation and temporal flicker.
[0009] Therefore, there is a need for enhanced systems and methods that can render perspective scenes for users without introducing visual artifacts, especially in the area the user is viewing. The present disclosure provides a solution through systems and methods that effectively address these challenges. Summary of the Invention
[0010] Embodiments described herein relate to an improved method for generating depth grids and using them to reproject captured images of a scene into a user's eye space for MR perspective generation. This disclosure balances the accuracy of the perceived perspective scene with the power, latency, and computational constraints of the system. This is achieved by utilizing eye tracking information to determine multiple three-dimensional locations (characterized by three-dimensional coordinates x, y, and z) in the scene that the user is looking at, and prioritizing these locations when generating a depth representation of the scene. In certain embodiments, the depth representation can be a depth grid with a single constant depth value corresponding to an object of interest derived from the user's gaze direction. For example, the depth grid can have a spherical outline (which can be a full sphere or a partial sphere), and each eye of the user can have its own depth grid of constant depth. The constant depth of the grid can vary based on the user's vergence and / or an estimate or prediction of the user's object of interest. For example, the MR system can use the user's head-mounted viewer's eye tracking module to determine the gaze of both eyes of the user. The MR system can then use the user's gaze information to determine the user's vergence location or object of interest. The distance between the user's left eye and the vergence position or object of interest can be used to generate a depth grid of constant depth for the left eye. Similarly, the distance between the user's right eye and the vergence position or object of interest can be used to generate another depth grid of constant depth for the right eye. The MR system can then use the user's respective depth grids for the left and right eyes to reproject images captured by the system's cameras to the user's left and right eyes.
[0011] The advantage of using depth meshes with constant depth is that these depth meshes are inherently stable and computationally inexpensive. Since depth meshes with constant depth do not have spatially varying depth (which typically includes highly inaccurate approximations between foreground and background objects), perspective scenes generated using depth meshes with constant depth are less prone to deformation and distortion. Furthermore, in a perspective scene, the region of interest that the user is viewing appears accurate because the depth mesh is generated based on the position of that region. Although objects that are closer or farther away than the region of interest may appear inaccurate, this inaccuracy has minimal negative impact on the overall perspective experience because the inaccurate parts of the scene are in the user's peripheral vision and are likely not of interest to the user. Therefore, reprojection using depth meshes with constant depth provides a practical solution to the aforementioned challenges.
[0012] In some aspects, the technology described herein relates to a method for displaying a scene to a user, the method comprising, by a computing device: receiving image data of a scene for display to the user, the image data comprising a left image for a left eye and a right image for a right eye; determining, by an eye tracking module, a gaze direction or eye convergence of the user; identifying an object in the scene on which the user is focusing using the gaze direction or the eye convergence; determining a left depth from the left eye to the object and a right depth from the right eye to the object; generating a left depth grid having a constant depth for the left eye based on the left depth; generating a right depth grid having a constant depth for the right eye based on the right depth; generating a left output image by projecting the left image onto the left depth grid; generating a right output image by projecting the right image onto the right depth grid; displaying the left output image for the left eye; and displaying the right output image for the right eye.
[0013] In some aspects, the technology described herein relates to a method in which a computing device is communicatively connected to a left camera and a right camera of a head-mounted device worn by a user, and in which a left image is acquired by the left camera and a right image is acquired by the right camera.
[0014] In some aspects, the technology described herein relates to a method in which an eye tracking module includes cameras pointed at a user's left and right eyes.
[0015] In some aspects, the technology described herein relates to a method in which the position of an object is determined based on the eye convergence of a user.
[0016] In some aspects, the techniques described herein relate to a method in which the position of an object is also determined based on scene information.
[0017] In some aspects, the technology described herein relates to a method in which identifying an object in a scene includes determining an intersection between a user's gaze direction and a depth of the scene.
[0018] In some aspects, the technology described herein relates to a method in which identifying an object in a scene includes determining an intersection between a user's gaze direction and a 3D model of the scene.
[0019] In some aspects, the techniques described herein relate to a method in which identifying objects in a scene is also based on a current usage context of a computing device.
[0020] In some aspects, the techniques described herein relate to a method in which the left and right depth grids are spherical.
[0021] In some aspects, the techniques described herein relate to a method in which the left and right depth grids are planar.
[0022] In some aspects, the techniques described herein relate to a method in which a constant depth of a left depth grid is different than a constant depth of a right depth grid.
[0023] In some aspects, the technology described herein relates to a method, further comprising: after generating the left depth grid and the right depth grid, using a second gaze direction or a second eye convergence of the user to identify a second object in the scene on which the user is focusing, wherein the second object is different from the above-mentioned object; determining a second left depth from the left eye to the second object and a second right depth from the right eye to the second object; generating a second left depth grid with a constant depth for the left eye based on the second left depth; generating a second right depth grid with a constant depth for the right eye based on the second right depth; generating a second left output image using the second left depth grid; generating a second right output image using the second right depth grid; displaying the second left output image for the left eye; and displaying the second right output image for the right eye.
[0024] In some aspects, the technology described herein relates to one or more computer-readable non-transitory storage media containing software that, when executed, is capable of: receiving image data of a scene around a user, the image data including a left image for a left eye and a right image for a right eye; determining a gaze direction or eye convergence of the user through an eye tracking module; using the gaze direction or the eye convergence to identify a region of interest in the scene; determining a left depth from the left eye to the region of interest and a right depth from the right eye to the region of interest; generating a left depth grid having a constant depth for the left eye based on the left depth; generating a right depth grid having a constant depth for the right eye based on the right depth; generating a left output image by projecting the left image onto the left depth grid; generating a right output image by projecting the right image onto the right depth grid; displaying the left output image for the left eye; and displaying the right output image for the right eye.
[0025] In some aspects, the techniques described herein relate to one or more computer-readable non-transitory storage media, wherein a location of a region of interest is determined based on eye convergence of a user.
[0026] In some aspects, the techniques described herein relate to one or more computer-readable non-transitory storage media, where the location of the region of interest is further determined based on scene information.
[0027] In some aspects, the techniques described herein relate to one or more computer-readable non-transitory storage media, wherein the left depth grid and the right depth grid are spherical.
[0028] In some aspects, the technology described herein relates to a system comprising: one or more processors; and one or more computer-readable non-transitory storage media, the one or more computer-readable non-transitory storage media being coupled to one or more of the one or more processors and storing instructions that, when executed by one or more of the one or more processors, are operable to cause the system to: receive image data of a scene surrounding a user, the image data comprising a left image for a left eye and a right image for a right eye; determine a gaze direction or eye convergence of the user through an eye tracking module; identify a region of interest in the scene using the gaze direction or the eye convergence; determine a left depth from the left eye to the region of interest and a right depth from the right eye to the region of interest; generate a left depth grid having a constant depth for the left eye based on the left depth; generate a right depth grid having a constant depth for the right eye based on the right depth; generate a left output image by projecting the left image onto the left depth grid; generate a right output image by projecting the right image onto the right depth grid; display the left output image for the left eye; and display the right output image for the right eye.
[0029] In some aspects, the technology described herein relates to a system in which the location of a region of interest is determined based on the vergence of a user's eyes.
[0030] In some aspects, the techniques described herein relate to a system in which the location of the region of interest is also determined based on scene information.
[0031] In some aspects, the techniques described herein relate to a system in which the left and right depth grids are spherical.
[0032] The embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited to these embodiments. A particular embodiment may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments according to the present disclosure are specifically disclosed in the accompanying claims relating to methods, storage media, and systems, wherein any feature mentioned in one claim category (e.g., method) may also be claimed in another claim category (e.g., system). Dependencies or back-references in the accompanying claims are selected for formal reasons only. However, any subject matter resulting from an intentional back-reference to any previous claim (particularly multiple dependent claims) may also be claimed, such that any combination of multiple claims and their multiple features is disclosed and may be claimed, regardless of the dependencies selected in the accompanying claims. Subject matter that may be claimed includes not only the combination of multiple features set forth in the accompanying claims, but also any other combination of multiple features in the claims, wherein each feature mentioned in a claim may be combined with any other feature or combination of features in the claims. Furthermore, any of the various embodiments and features described or depicted herein may be claimed in a separate claim and / or in any combination with any embodiment or feature described or depicted herein or in any combination with any feature in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is an illustrative system for presenting a scene to a user according to the disclosed embodiments.
[0034] Figure 2A is a schematic representation of a system for displaying a scene to a user according to the disclosed embodiments.
[0035] Figure 2B is another representation of a system for displaying a scene to a user according to the disclosed embodiments.
[0036] Figure 3A is a schematic representation of a variable depth mesh and objects located in a user environment according to the disclosed embodiments.
[0037] Figure 3B is a schematic diagram illustrating a constant depth grid, a variable depth grid, and a system for displaying a scene to a user according to the disclosed embodiments.
[0038] Figure 3C is a schematic diagram illustrating projecting an image of an object onto a constant depth grid according to the disclosed embodiments.
[0039] Figure 3D is a schematic diagram illustrating projecting images of multiple objects onto a constant depth grid and a variable depth grid according to the disclosed embodiment.
[0040] Figure 3E is a schematic diagram illustrating projecting images of multiple objects onto a constant depth grid according to the disclosed embodiment.
[0041] Figure 4 is a diagram listing several methods for displaying scenes to a user based on the user's movement and location according to the disclosed embodiments.
[0042] Figure 5 is an example method of displaying a scene to a user by projecting an image of an object onto a constant depth grid according to the disclosed embodiments.
[0043] Figure 6 is another example method of displaying a scene to a user by projecting at least some portions of an image of an object onto a constant depth grid according to the disclosed embodiments.
[0044] Figure 7 is another example method according to the disclosed embodiments of displaying a scene to a user by projecting objects near the user onto a constant depth grid. DETAILED DESCRIPTION
[0045] In the following description, for the purpose of explanation, numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent that embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring the description of the present disclosure.
[0046] The text of this disclosure, in conjunction with the accompanying drawings, is intended to present the algorithms required to program a computer to implement various embodiments in a straightforward manner, with the same level of detail that persons skilled in the art would typically use to communicate with one another about the functions to be programmed, inputs, transformations, outputs, and other aspects of programming. In other words, the level of detail set forth in this disclosure is the same level of detail that persons skilled in the art would typically use to communicate with one another about the structure and functions of an algorithm or program to be programmed to implement the embodiments of this disclosure.
[0047] Various embodiments may be described in this disclosure to illustrate various aspects. Other embodiments may also be utilized and structural, logical, software, electrical, and other changes may be made to these other embodiments without departing from the scope of the specifically described embodiments. Various modifications and alterations are possible and contemplated. Some features may be described with reference to one or more embodiments or figures, but these features are not limited to use in the one or more embodiments or figures described with reference to these features. Therefore, this disclosure is neither a literal description of all embodiments nor an enumeration of features that must be present in all embodiments.
[0048] Devices described as being in communication with each other need not necessarily be in continuous communication with each other unless expressly stated otherwise. In addition, devices that are in communication with each other may communicate directly or indirectly through one or more logical or physical intermediaries.
[0049] The description of an embodiment in which multiple components communicate with each other does not mean that all of these components are required. Optional components may be described to illustrate various possible embodiments and to more fully illustrate one or more aspects of the present disclosure.
[0050] Similarly, although processing steps, method steps or algorithms can be described in sequence, unless there is a clear statement to the contrary, such processes, methods and algorithms can generally be configured to work in different orders. Any sequence or order of the steps described in this disclosure is not a necessary sequence or order. The steps of the described process can be performed in any actual order. In addition, some steps can be performed simultaneously. The description of the process in the accompanying drawings does not exclude variations and modifications, does not mean that the process or any step thereof is necessary, and does not mean that the process shown is preferred. Each embodiment can describe these steps once, but does not have to be performed only once. Some steps can be omitted in some embodiments or some events, or some steps can be performed more than once in each embodiment or event. When describing a single device or article, more than one device or article can be used to replace a single device or article. In the case of describing more than one device or article, a single device or article can be used to replace more than one device or article.
[0051] The functions or features of a device may alternatively be implemented by one or more other devices that are not explicitly described as having such functions or features. Therefore, other embodiments need not include the device itself. For clarity, the techniques and mechanisms described or referenced herein will sometimes be described in the singular. However, it should be noted that, unless otherwise stated, the embodiments include multiple iterations of the technology or multiple manifestations of the mechanism. The process descriptions or blocks in the accompanying drawings should be understood to represent modules, segments or code portions that include one or more executable instructions for implementing specific logical functions or steps in the process. Alternative implementations are included within the scope of the embodiments of the present disclosure, wherein, for example, depending on the functions involved, the functions can be performed in the order shown or discussed, including substantially simultaneously or in reverse order.
[0052] System Overview
[0053] Embodiments presented herein relate to systems and methods designed to render see-through scenes to a user using various display options, with one example being a virtual reality headset. These see-through scenes can include various scenes from the user's physical environment, such as a room, a house, a playground, and a landscape.
[0054] Furthermore, the scope of display devices is not limited to virtual reality headsets; it extends to a variety of alternatives. These alternatives include video screens, smartphones, glasses, smart augmented reality glasses, camera viewfinders, telescopes, binoculars, microscopes, and similar devices.
[0055] As previously described, perspective rendering is achieved by capturing an image of the scene using a suitable camera of the head-mounted view (e.g., an externally facing RGB camera or a monochrome camera of the head-mounted view for capturing images for perspective generation), reprojecting the captured image onto a depth grid of the user's environment to generate perspective images for both eyes of the user, and displaying the perspective images to both eyes of the user. In some cases, one perspective image may be presented to the user's left eye, while another perspective image may be presented to the user's right eye. In various embodiments, the various steps of rendering the scene are implemented by a system including a computing device, an image capture module, an eye tracking module, a depth estimation module, and a display module.
[0056] For example, Figure 1An example system 100 for rendering a scene to a user based on acquired image data and a determined depth grid is shown. In various embodiments, the system 100 can perform one or more steps of one or more methods described or illustrated herein. The system 100 may include software instructions for performing one or more steps of the methods described or illustrated herein. Additionally, various other instructions may provide various other functionality of the system 100, as described or illustrated herein. Various embodiments include one or more portions of the system 100. The system 100 may include one or more computing systems. Herein, references to a computer system may include computing devices, and vice versa, where appropriate. Additionally, references to a computer system may include one or more computer systems, where appropriate.
[0057] The present disclosure contemplates any suitable number of computer systems that may be included in system 100. The present disclosure contemplates system 100 taking any suitable physical form. By way of example and not limitation, system 100 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or a system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a network of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, a game console, or a combination of two or more of these systems. Where appropriate, system 100 may include one or more computer systems, may be single or distributed, may span multiple locations, may span multiple machines, may span multiple data centers, or may be located in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, the system 100 can perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, the system 100 can perform one or more steps of one or more methods described or illustrated herein in real time or in a batch mode. Where appropriate, the system 100 can perform one or more steps of one or more methods described or illustrated herein at different times or at different locations.
[0058] In various embodiments, system 100 includes a computing device 101 that includes a processor 102, memory 104, storage 106, an input / output (I / O) interface 108, a communication interface 110, and a bus 112. Additionally, system 100 includes an image acquisition module 120, an eye tracking module 130, a depth estimation module 140, a display module 150, and optionally a lighting module 160 and a motion acquisition module 170. Although this disclosure describes and illustrates a particular system 100 having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable system having any suitable number of any suitable components in any suitable arrangement.
[0059] In various embodiments, processor 102 includes hardware for executing instructions, such as those that constitute a computer program. By way of example and not limitation, to execute instructions, processor 102 may retrieve (or fetch) the instructions from internal registers, an internal cache, memory 104, or storage 106; decode and execute the instructions; and then write one or more results to the internal registers, internal cache, memory 104, or storage 106. In some embodiments, processor 102 may include one or more internal caches for data, instructions, or addresses. Where appropriate, the present disclosure contemplates processor 102 including any suitable number of any suitable internal caches. By way of example and not limitation, processor 102 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The instructions in the instruction caches may be copies of instructions in memory 104 or storage 106, and the instruction caches may speed up retrieval of these instructions by processor 102. The data in the data cache may be a copy of data in memory 104 or storage 106 for instruction operations executed at processor 102; may be the result of a previous instruction executed at processor 102, which is used for subsequent instructions executed at processor 102 to access or write to memory 104 or storage 106; or may be other suitable data. The data cache may accelerate read or write operations of processor 102. The TLB may accelerate virtual address translation by processor 102. In certain embodiments, processor 102 may include one or more internal registers for data, instructions, or addresses. Where appropriate, the present disclosure contemplates processors 102 including any suitable number of any suitable internal registers. Where appropriate, processor 102 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include one or more processors 102. Although the present disclosure describes and illustrates a particular processor, the present disclosure contemplates any suitable processor.
[0060] In certain embodiments, memory 104 includes main memory for storing instructions for processor 102 to execute or data for processor 102 to operate on. By way of example and not limitation, system 100 may load instructions from storage device 106 or another source (e.g., another system 100) into memory 104. Processor 102 may then load these instructions from memory 104 into internal registers or an internal cache. To execute these instructions, processor 102 may retrieve these instructions from the internal registers or internal cache and decode them. During or after executing these instructions, processor 102 may write one or more results (which may be intermediate results or final results) to the internal registers or internal cache. Processor 102 may then write one or more of these results to memory 104. In certain embodiments, processor 102 executes instructions from one or more internal registers or internal caches, or from memory 104 (rather than from storage device 106 or elsewhere), and operates on data from one or more internal registers or internal caches, or from memory 104 (rather than from storage device 106 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple the processor 102 to the memory 104. As described below, the bus 112 may include one or more memory buses. In particular embodiments, one or more memory management units (MMUs) are located between the processor 102 and the memory 104 and facilitate access to the memory 104 requested by the processor 102. In particular embodiments, the memory 104 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port RAM or multi-port RAM. The present disclosure contemplates any suitable RAM. Where appropriate, the memory 104 may include one or more memories 104. Although the present disclosure describes and illustrates specific memories, the present disclosure contemplates any suitable memory.
[0061] In certain embodiments, storage device 106 comprises a mass storage device for data or instructions. By way of example and not limitation, storage device 106 may comprise a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these storage devices. Where appropriate, storage device 106 may comprise removable media or non-removable (or fixed) media. Where appropriate, storage device 106 may be located internally or externally to system 100. In certain embodiments, storage device 106 is a non-volatile solid-state memory. In certain embodiments, storage device 106 comprises read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), flash memory, or a combination of two or more of these ROMs. This disclosure contemplates mass storage 106 taking any suitable physical form. Where appropriate, storage 106 may include one or more storage control units that facilitate communication between processor 102 and storage 106. Where appropriate, storage 106 may include one or more storage devices 106. Although this disclosure describes and illustrates particular storage devices, this disclosure contemplates any suitable storage devices.
[0062] In certain embodiments, the I / O interface 108 includes hardware, software, or both that provide one or more interfaces for communication between the system 100 and one or more I / O devices. Where appropriate, the system 100 may include one or more of these I / O devices. One or more of these I / O devices may enable communication between an individual and the system 100. By way of example and not limitation, an I / O device may include a keyboard, a keypad, a microphone, a monitor, a mouse, a printer, a scanner, a speaker, a still camera, a stylus, a tablet computer, a touch screen, a trackball, a video camera, another suitable I / O device, or a combination of two or more of these I / O devices. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 108 for these I / O devices. Where appropriate, the I / O interface 108 may include one or more device or software drivers that enable the processor 102 to drive one or more of these I / O devices. Where appropriate, I / O interface 108 may include one or more I / O interfaces 108. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.
[0063] In certain embodiments, the communication interface 110 includes hardware, software, or both that provides one or more interfaces for communication (e.g., packet-based communication) between the system 100 and any other device interacting with the system 100 via one or more networks. By way of example and not limitation, the communication interface 110 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network (e.g., a WI-FI network). The present disclosure contemplates any suitable network and any suitable communication interface 110 for the network. By way of example and not limitation, the system 100 may communicate with an ad hoc network; a personal area network (PAN); a local area network (LAN); a wide area network (WAN); a metropolitan area network (MAN); or one or more portions of the Internet; or a combination of two or more of these networks. One or more portions of one or more of these networks may be wired or wireless. For example, the system 100 may communicate with a wireless PAN (WPAN) (e.g., a Bluetooth WPAN), a Wi-Fi network, a WI-MAX network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), or other suitable wireless networks, or a combination of two or more of these networks. Where appropriate, the system 100 may include any suitable communication interface 110 for any of these networks. Where appropriate, the communication interface 110 may include one or more communication interfaces 110. Although this disclosure describes and illustrates particular communication interfaces, this disclosure contemplates any suitable communication interface.
[0064] In particular embodiments, bus 112 includes hardware, software, or both that couples the various components of system 100. By way of example and not limitation, bus 112 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a low-pin-count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, another suitable bus, or a combination of two or more of these busses. Where appropriate, bus 112 may include one or more buses 112. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0065] As used herein, the one or more computer-readable non-transitory storage media may include, where appropriate, one or more semiconductor-based integrated circuits (ICs) or other ICs (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disc drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or secure digital drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, the computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0066] In various embodiments, the system 100 includes an image acquisition module 120, which includes one or more image acquisition devices. These devices can be composed of a variety of suitable cameras, including those designed to capture visible light, infrared light, or ultraviolet light. Camera options can use complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) sensors, which have different sizes and resolutions. These sensors include full-frame sensors, advanced photographic system-classic (APS-C) sensors, and compact sensors, with resolutions ranging from several million pixels to tens of millions of pixels, or even more than 50 million pixels. In addition, the image acquisition device can be equipped with various lens systems, such as a zoom lens, a wide-angle lens, a fisheye lens, a telephoto lens, a macro lens, a tilt-shift lens, or any other suitable lens. The image acquisition module 120 may also include additional components such as a flash (e.g., a light emitting diode (LED) flash), an independent power source (e.g., a battery) for operating the image acquisition device, and a local data storage device for on-site image data storage.
[0067] In certain embodiments, the image acquisition module 120 may include one or more image acquisition devices for acquiring images to be re-projected to the viewpoints of the user's eyes, thereby generating a perspective image. For example, a first image acquisition device may be used to acquire a first image of a scene, and a second image acquisition device may be used to acquire a second image of the same scene. These image acquisition devices may be positioned at a predetermined distance from each other. For example, in the context of a virtual reality (VR) head-mounted viewer, the first image acquisition device may be located near one of the user's eyes (e.g., the left eye) and referred to herein as a left camera, while the second image acquisition device may be located near the user's other eye (e.g., the right eye) and referred to herein as a right camera. The image acquired by the left camera may be re-projected to generate a perspective image for the user's left eye, while the image acquired by the right camera may be re-projected to generate a perspective image for the user's right eye. The distance between the two devices is typically approximately the same as the distance between the user's eyes, typically in the range of about 55 mm to about 80 mm.
[0068] In some embodiments, each image acquisition device may include one or more cameras, which are generally referred to as sensors. For example, the image acquisition device within the image acquisition module 120 may include a main sensor, such as a high-resolution sensor (e.g., an 8- to 30-megapixel sensor) having a digital sensor array of pixel sensors. Each pixel sensor can be designed to have a sufficiently large size (e.g., a few microns) to capture a sufficient amount of light from the aperture of the camera. In addition, the image acquisition device for perspective or MR generation may include an ultra-wide sensor, a telephoto sensor, or any other suitable sensor, and these sensors can be configured to acquire color image data or monochrome image data.
[0069] The eye tracking module 130 of the system 100 can be configured to monitor the movement (including rotation) of the user's eyes. The MR system 100 can use the eye tracking module 130 to determine the convergence position of the user's eyes and / or the area or object of interest to the user. The present disclosure is not limited to any particular type of eye tracking module 130. For example, the system 100 can include an eye tracking camera that is customized to detect eye movements, thereby enabling the determination of the visual axes of the left and right eyes and the user's gaze direction. Various devices for eye tracking can be used, including trackers for video-based tracking (e.g., cameras), infrared trackers, or any other suitable eye tracking sensors (e.g., electrooculogram sensors). Once the image acquisition device identifies an object within a scene, it can be configured to adjust its focus to focus on an area around the identified object of the scene.
[0070] The depth estimation module 140 of the system 100 is configured to estimate depth information of the user's physical environment and generate a corresponding depth grid. The depth estimation module 140 can use any suitable technology for estimating depth. For example, the stereo depth estimation module 140 may include a computing device that executes programming instructions to calculate the distance to an object using triangulation. Additionally or alternatively, the depth estimation module 140 can utilize various depth measurement devices, such as a stereo camera (e.g., a stereo camera uses two or more cameras to capture images of a scene from slightly different angles, and the parallax between corresponding elements in the captured images can be used to estimate depth), a structured light scanner (e.g., a device that projects a structured light pattern such as a grid or stripes onto an object or environment and calculates depth based on the deformation of the structured pattern), a laser radar (LiDAR), a time-of-flight (TOF) sensor, an ultrasonic sensor, or any other suitable sensor that can be used to estimate depth (e.g., a mobile camera using photogrammetry).
[0071] The display module 150 of the system 100 is configured to display a scene to the user. The display module 150 may include one or more displays. For example, when the display module 150 is part of a virtual head-mounted view device, the left display may be configured to display the scene to the user's left eye using image data captured by a left camera, while the right display may be configured to display the scene to the user's right eye using image data captured by a right camera, thereby simulating depth perception and providing the user with three-dimensional rendering. In some cases, the left and right displays may include corresponding left and right lenses to optimize the field of view and simulate natural vision. The virtual head-mounted view device displays may use technologies such as liquid crystal displays (LCDs) or organic light-emitting diodes (OLEDs) to produce high-quality visual effects. In addition to displaying image data from the cameras, these displays may also display virtual objects within the scene, annotations to scene objects, or transformations performed on objects (e.g., recoloring or reshaping) and display modified versions.
[0072] Additionally, in some embodiments, the system 100 can include a lighting module 160 configured to enhance the illumination of various objects within a scene viewed by a user. The lighting module can enhance the ability of the system 100 to capture images for perspective generation, depth estimation, and / or tracking. The lighting module 160 can include one or more lighting devices, which can include various options for emitting light to illuminate an object, including, but not limited to, light-emitting diodes (LEDs), infrared light emitters (e.g., infrared LEDs), fluorescent light sources, compact fluorescent light sources, incandescent light sources, halogen light sources, lasers, organic LEDs, black light sources, and ultraviolet light sources.
[0073] Furthermore, in some embodiments, the system 100 may include a motion capture module 170 (e.g., an accelerometer or gyroscope). This module is designed to capture the movements of the image capture module 120, the depth estimation module 140, and the lighting module 160. In the event that the image capture module 120, the depth estimation module 140, and the lighting module 160 are integrated into a VR headset, the motion capture module 170 may be configured to record the translational and rotational displacements of the VR headset. This includes instances where the user rotates or moves their head.
[0074] Data regarding these displacements has a variety of uses. This data can be used to recalibrate the depth information collected by depth estimation module 140, to reacquire new depth data using depth estimation module 140, or to acquire new image data using image acquisition module 120. Furthermore, the movement of the user's head can indicate a shift in the user's focus to a new object within the observed scene. This change in focus can, in turn, trigger further movement and adjustment in the image acquisition device of image acquisition module 120 and the light source of light emission module 160.
[0075] In some scenarios, the motion capture module 170 may be configured not only to detect movement of the user's head, but also to monitor the movement of one or more objects within the user's environment. When such object movement is detected, the system 100 may determine that new image data needs to be captured by the image capture module 120 or that a new depth grid needs to be generated using the depth estimation module 140. For example, if an object (e.g., a pet) surrounding the user within a scene moves within the scene, new image and depth data may be collected.
[0076] In various embodiments, the system 100 may take the form of a virtual reality (VR) head-mounted viewer. For example, Figure 2A and Figure 2BAn exemplary embodiment of system 200 is shown, which represents an implementation of system 100. System 200 includes an image acquisition module 220 that houses a first image acquisition device 220L (e.g., a left camera) and a second image acquisition device 220R (e.g., a right camera), a depth estimation module 240, a lighting module 260, and a motion estimation module 270. The modules of system 200 can be similar in structure and function to corresponding modules of system 100.
[0077] exist Figure 2A In the schematic representation of system 200 in FIG, a left camera 220L and a right camera 220R are positioned proximate to the user's respective left eye 205L and right eye 205R as viewed in cross section H. A left line of sight 206L (also referred to as a left viewing distance or left ray 206L), representing the orientation of left eye 205L, is shown passing through left camera 220L. Figure 2A The angle θ formed by the left line of sight 206L and the normal direction N is shown, which determines the horizontal orientation of the left eye 205L. Another orientation angle can indicate the vertical deviation of the left line of sight 206L from the normal direction N in a plane perpendicular to plane H. Similarly, the right line of sight 206R represents the orientation of the right eye 205R. Although the right line of sight 206R does not directly pass through the right camera 220R, it is close enough. This proximity ensures that the image captured by the right camera 220R is very similar to the image that the right eye 205R would naturally observe without a virtual reality headset.
[0078] In addition, if Figure 2A As shown, system 100 may include an eye tracking module 230, which may include a left eye tracking device 230L and a right eye tracking device 230R. Eye tracking module 230 may be similar in function or structure to eye tracking module 130. These devices may include suitable cameras (as described above). In an example embodiment, left eye tracking device 230L may be configured to determine the orientation of left eye 205L, while right eye tracking device 230R may be configured to determine the orientation of right eye 205R.
[0079] Method Overview
[0080] In various embodiments described herein, a method for presenting a see-through scene to a user involves generating a depth grid and acquiring image data of the see-through scene. The left image and the right image collected from the left camera (e.g., the left camera 220L of the system 200) and the right camera (e.g., the right camera 220R of the system 200) are then re-projected to the left eye 205L and the right eye 205R, respectively, through the depth grid.
[0081] Figure 3AAn example scene 300 is shown in which a depth grid having variable depth (also referred to herein as a variable depth grid), represented by curve 343, is shown along with images of objects 318 and 317. In addition, Figure 3A An area 310 representing the horizon is shown, which is assumed to maintain a constant depth. Figure 3A is a top-down, two-dimensional view of the three-dimensional scene 300. The depth grid 343 can be a continuous surface that is deformed based on the positions of the objects 317 and 318. For example, the depth estimation module 240 can estimate depth information in the scene, which includes the depths of the objects 317 and 318. The depth grid 343 can be deformed so that its contour generally matches the estimated depth in the scene, so that the deformed depth grid 343 provides a representation of the depth of the scene. The image captured using the camera (e.g., 220L or 220R) is re-projected to the corresponding eye (e.g., 205L or 205R) of the user through the depth grid 343 to generate a perspective image.
[0082] Artifacts occur when the depth mesh does not accurately represent the physical scene depth. For example, Figure 3A Depth grid 343 in
[0045] can accurately reflect the depths of objects 317 and 318. However, the region of depth grid 343 between objects 317 and 318 does not accurately reflect scene 300. This may be due to poor or inaccurate depth estimation and / or one of the drawbacks of using a continuous grid to represent scene depth. Because the region of depth grid 343 between objects 317 and 318 is inaccurate and varies greatly, the corresponding region in the re-projected perspective image will often exhibit distortion and deformation artifacts.
[0083] To address the issues associated with the artifacts described above, certain embodiments described herein use a depth grid having a constant depth that is determined based on the object that the user is focusing on. Figure 3B Shows the use of a depth grid with constant depth instead of using a reference Figure 3A An example of a variable depth grid 343 is depicted. Figure 3B The schematic diagram in shows that the user is focusing on object 318 by displaying the user's overall gaze direction 308, which is determined by an eye tracking module that is similar in structure or function to (or identical to) eye tracking module 130. The orientations of left eye 305L and right eye 305R are indicated by corresponding left and right gaze lines 306L, 306R, which converge on object 318, confirming the user's focus.
[0084] Once the system has determined the vergence of the user's gazes 306L, 306R, it can calculate the distance between the vergence position and each of the user's eyes. In embodiments where the vergence position is used to determine the desired depth of the constant depth grid, depth information of the scene is not required because the vergence position can be calculated based solely on eye tracking data. Figure 3B In the example shown, the user's convergence is at object 318. The system can use any known technique to determine the distance between object 318 and the user's eyes. For example, if the distance D between the left iris of the left eye 305L and the right iris of the right eye 305R is LR is known (such as Figure 3B depicted and can be determined using an eye tracking module), and the angle θ is determined by the eye tracking module L and θ R , the distance from the left eye 305L to the object 318 along the left line of sight 306L can be calculated, thereby providing the left depth. Similarly, knowing θ L ,θ R , and D LR It helps to calculate the distance from the right eye 305R to the object 318 along the right line of sight 306R, thereby establishing the right depth.
[0085] In another embodiment, the expected depth of the constant depth grid can be determined with the help of additional scene information, such as scene depth and / or contextual data. Such scene information can be used as an alternative or in addition to the aforementioned vergence estimation to determine the user's region of interest. For example, if the scene depth is known (e.g., based on depth measurements using depth estimation module 240), the intersection of the user's gaze and the scene depth can be used as the region of interest. Similarly, in embodiments where the system has a stored 3D model of the user's environment, the intersection between the user's gaze and the 3D model can be used as the region of interest. This 3D model can be generated based on any suitable 3D reconstruction technique. As another example, based on the current usage context, the system can predict the user's likely region of interest. For example, since the 3D location of the MR content is known to the system, the system can use the 3D location of the specific MR content the user is currently engaging with as the expected region of interest. Any combination of these examples of using scene information to determine the user's region of interest can be used in conjunction with vergence estimation to improve the system's overall prediction of the user's region of interest. Once the system identifies the region of interest, it can calculate the depth of that region from each of the user's eyes.
[0086] Once the left and right depths are determined, a left constant depth grid and a right constant depth grid can be generated. For example, Figure 3BA left depth grid (also referred to as a left constant depth grid) represented by arc 342 and a right depth grid (also referred to as a right constant depth grid) represented by arc 341 are shown. The left depth grid 342 has a constant left depth corresponding to the distance between the user's left eye 305L and the object 318, while the right depth grid 341 has a constant right depth corresponding to the distance between the user's right eye 305R and the object 318. Figure 3B As shown, since the right eye 305R is at a greater distance from the object 318 , the right depth is greater than the left depth, as shown by the difference in radius between arc 342 and arc 341 .
[0087] The constant depth grids 342, 341 can then be used to generate perspective images. Cameras 320L and 320R are configured to capture image data of a scene 300 including an object 318. In an exemplary embodiment, camera 320L captures a left image for a left eye 305L, and camera 320R captures a right image for a right eye 305R. Figure 3B As can be seen in the figure, the left camera 320L and the left eye 305L are not in exactly the same position, and the same is true for the right camera 320R and the right eye 305R. Due to this difference, the object 318 (and generally the scene 300) will appear slightly different from the perspective of the cameras 320L, 320R and the user's eyes 305L, 305R. As shown in the figure, the line of sight 307L, 307R from the cameras 320L, 320R to the object 318 is different from the line of sight 306L, 306R from the user's eyes 305L, 305R to the object 318. Therefore, the images captured by the cameras 320L, 320R need to be re-projected to the eyes 305L, 305R, respectively, so that the perspective image of the scene 300 has the correct viewer's perspective. Once the left and right images are collected, they are re-projected using a corresponding left depth grid 342 with a constant left depth and a right depth grid 341 with a constant right depth, respectively. Image projection involves transforming image blocks (e.g., stretching, tilting, resizing, or shrinking blocks) so that the transformed image represents a projection of the image captured by the camera onto the depth grid surface and is rendered for the user's eyes. For example, the left image is projected onto the left depth grid 342, and the right image is projected onto the right depth grid 341. The left projected image is then displayed for the user's left eye 305L, which may be referred to as a left perspective image. Similarly, the right projected image is displayed for the user's right eye 305R, which may be referred to as a right perspective image. In this example, because the constant depth grids 342, 341 reflect the correct depth of object 318, object 318 appears correct in the perspective image. Because the user is focusing on object 318, it is important for object 318 to appear accurate and have minimal artifacts.
[0088] While object 318 may appear accurate in the perspective image, other objects at different depths in scene 300 may not be rendered accurately. For example, peripheral object 317 is much farther from the user than object 318. Therefore, when object 317 is reprojected using constant depth grids 342, 341 (which in this example is optimized for object 318), the reprojection of object 317 may appear inaccurate (e.g., it may appear distorted or deformed). This is particularly true in the following examples: Figure 3C As shown in Figure 3C In this example, a left depth grid 342 and a right depth grid 341 of constant depth are used to reproject peripheral objects such as object 317. Image projection 316L represents the projection of object 317 onto left depth grid 342, and image projection 316R represents the projection of object 317 onto right depth grid 341. Because constant depth grids 342 and 341 are optimized for object 318, image projections 316L and 316R are misaligned and do not accurately reflect object 317. The end result may be that the object is presented to the user as image 319 at the point where left and right sight lines 306L and 306R intersect, or artifacts such as distortion and blurring may occur. However, because the user in this example is focusing on object 318 rather than object 317, any inaccurate visual representation of object 317 may not be noticed or have minimal impact on the user's overall experience. In certain embodiments, the system may also mitigate inaccurate visual representations of object 317 by applying one or more filters, such as by blurring areas in the user's periphery, so that inaccuracies and distortion artifacts are less noticeable.
[0089] In certain embodiments, the MR system can dynamically adjust the constant depth grids 341 and 342 as needed. For example, the MR system can periodically repeat the following process: determining the user's gaze direction; identifying regions of interest and their corresponding depths; using the depths to reconfigure the constant depth grids for the left and right eyes; and reprojecting the acquired images using the constant depth grids to generate perspective images. In certain embodiments, this process can be repeated at a predetermined rhythm (e.g., the process can be repeated once per frame or once every more than one frame). In other embodiments, the constant depth grids can be reconfigured when the system determines that the user's gaze has shifted away from an object or if the object has moved. For example, if the user's gaze shifts from object 318 to object 317, the system can calculate the depth of object 317 for the left constant depth grid and another depth for object 317 for the right constant depth grid. Now that the user's gaze has shifted, object 318 becomes a peripheral object and may appear distorted or deformed in the perspective image generated using the new constant depth grid.
[0090] Various alternative hybrid approaches can be employed that combine projecting certain objects or parts of the scene onto a constant depth grid while projecting other objects or parts of the scene onto a variable estimated depth grid. Figure 3D An example of such an embodiment is shown in FIG. Here, the user focuses on object 318, causing constant depth grids 341, 342 to be configured based on the depth of object 318. Additional object 315 is within the user's viewing cone 333 (e.g., a predefined limited field of view). Since object 315 is within viewing cone 333, object 315 is reprojected to the user's viewpoint using left constant depth grid 342 and right constant depth grid 341. Figure 3D As depicted, the left projection of object 315 is represented by 314L along the left line of sight 306L, and the right projection of object 315 is represented by 314R along the right line of sight 306R. Objects located outside of the viewing cone 333 (e.g., object 317) can be reprojected to the user's viewpoint via a variable depth grid 343. An advantage of this is that the variable depth grid 343 can be more accurate for object 317 than the constant depth grids 341, 342. The viewing cone 333 can be determined using any suitable method. In an exemplary embodiment, the viewing cone 333 can have a vertex located between the user's left eye 305L and right eye 305R (e.g., at a midpoint between the user's eyes) and can have a cone axis aligned with the gaze direction 308. Additionally, as shown Figure 3D As shown, the aperture can take any suitable value, which can be, for example, in the range of 5 degrees to 100 degrees, or can have any other suitable value. In some cases, selecting the aperture can include determining the maximum feature size of an object observed in the left image and the right image, and selecting the diameter of the bottom of the viewing cone to be at least the maximum feature size of the object. Based on such a selection, the aperture can then be calculated as two times the inverse tangent of the ratio between the diameter and the depth of the object, where the depth of the object is estimated by the depth estimation module along the gaze direction 308.
[0091] Figure 3E An alternative mixing method is shown. Similar to reference Figure 3D The described embodiment uses a constant depth grid generated based on the user's region of interest (e.g., constant depth grids 341 and 342, which are generated based on the user's region of interest). Figure 3D but for simplicity in Figure 3E333). However, unlike previous embodiments, a large default constant depth grid 344 is used instead of a variable depth grid 343 to reproject objects located outside the viewing cone 333. For example, objects 317 and 313 located outside the user's viewing cone 333 can be reprojected to the user's viewpoint using the default constant depth grid 344. The default constant depth grid 344 can be particularly suitable for reprojecting background scenes. The default constant depth grid 344 can have a greater depth than either of the constant depth grids 341, 342. The depth of the larger constant depth grid 344 can be a predetermined default value or calculated based on an average or approximate value of the depths of objects outside the viewing cone 333. In certain embodiments, a single default constant depth grid 344 can be utilized to reproject both the left and right images of objects 317 and 313.
[0092] In various ways, Figures 3A to 3E The methods shown are combined to create a method for rendering objects within a three-dimensional scene for a user. For example, images of some objects within a first viewing frustum can be projected onto a first left depth grid having a first constant left depth and a first right depth grid having a first constant right depth, while images of other objects within a second viewing frustum but outside the first viewing frustum can be projected onto second left and second right depth grids having corresponding second constant left and second constant right depths. In some scenarios, the same constant depth grid can be used to project both the left and right images. In other scenarios, images of objects near the user can be projected onto left and right grids of constant depth, while images of objects farther away from the user can be projected onto a larger constant depth grid or a variable estimated depth grid for that scene.
[0093] The constant depth grid reprojection technique may not be a suitable solution for all scenarios. For example, reprojection based on a constant depth grid is suitable for situations where the scene depth is relatively constant, such as when the user is sitting. However, reprojection based on a constant depth grid is less suitable for situations where the depth of the environment varies significantly and the user may be viewing various objects at different depths (e.g., when the user is walking). Therefore, in certain embodiments, the MR system can adaptively select a specific type of depth grid for reprojection based on the user's current context. The MR system can determine the current usage context by analyzing sensor data collected by one or more types of sensors and technologies (e.g., accelerometers, gyroscopes, inertial measurement units, cameras, depth sensors, positioning and tracking technologies, etc.). For example, the MR system can use this sensor data to monitor the user's movement and predict whether the user is likely to be stationary and fixed in a specific direction (i.e., the user exhibits no lateral or rotational movement), stationary but looking around, not in a specific direction (i.e., the user exhibits rotational movement but no lateral movement), or moving and looking around (i.e., the user exhibits both rotational movement and lateral movement). The MR system can also use its understanding of the user's current state in the MR application to predict the user's usage context. For example, if a user is watching a virtual television (TV) in MR, the user is likely stationary. As another example, if a user is using a navigation MR application, the MR system may infer that the user is likely walking outdoors. Those skilled in the art will recognize that the MR system can use any combination of sensor data, tracking technology, and MR application state to determine the user's current usage context.
[0094] The MR system can adaptively select different types of depth grids for perspective generation based on the user's usage context. Figure 4 Three different methods (methods 1 to 3) are schematically shown, which can be customized based on the user's actions or usage context. For example, when user A is stationary and viewing a scene with approximately uniform depth in roughly the same direction (e.g., watching TV), the MR system can select method 1, which may involve reprojecting using a planar constant depth grid. In this use case, using a planar constant depth grid may be advantageous because a planar grid is a good approximation of a substantially uniform scene depth (e.g., the side of a room where a TV is placed is typically planar). The planar constant depth grid can be positioned in the user's viewing direction and oriented to match the scene. Similar to the spherical constant depth grid described previously, the depth of the planar constant depth grid can be adjusted based on the user's area of interest.
[0095] In another use case where the user exhibits rotational movement but is stationary (e.g. Figure 4(where user B stands or sits at the same location but can look in different directions), the MR system may choose method 2 to generate perspective images. For example, method 2 may involve using a spherical constant depth grid, similar to reference
[15] . Figures 3A to 3E The technology described.
[0096] In yet another use case, a user may exhibit both rotational and lateral movement (e.g., Figure 4 ( User C in the figure is walking). Especially when the user moves in this way in an environment with significant depth variations (e.g., walking outdoors), a constant depth grid may not provide optimal results. Therefore, when the user exhibits both rotational and lateral movement, the MR system may choose to use a variable depth grid instead of a constant depth grid, as the variable depth grid will be better able to adapt to complex scenes.
[0097] Instead of or in addition to considering the user's movement, the MR system can consider the approximate distance between the user and the scene of interest when selecting the type of depth grid to use. As previously described, the MR system can use eye tracking / vergence information and / or information about the scene (e.g., depth information collected using a depth sensor, contextual information, etc.) to determine the approximate distance of the scene of interest from the user. When the MR system determines that the scene of interest is within a threshold distance considered close to the user (e.g., within a threshold of 1 meter, 1.5 meters, or 2 meters), the MR system can choose to use a constant depth grid for perspective generation. The specific type of constant depth grid used (e.g., spherical or planar) also depends on whether the user is likely to exhibit rotational movement. For example, when the user is viewing something nearby and is unlikely to exhibit rotational movement (e.g., the user is viewing a display or reading content), the MR system can choose to use a planar constant depth grid for reprojection. On the other hand, when the user is viewing something nearby and is likely to exhibit rotational movement (e.g., the user is not viewing any specific content), the MR system can choose to use a spherical constant depth grid. In scenarios where the user is viewing a scene that is far away (e.g., beyond a predetermined threshold, such as 2 meters, 3 meters, 5 meters, or 10 meters), the MR system may choose to use a variable depth grid instead of a constant depth grid to perform perspective-generated re-projection. In certain embodiments, the MR system may also distinguish between intermediate-distance scenes and far-distance scenes. For example, if the user is viewing a scene within a predetermined intermediate range (e.g., between 2 meters and 4 meters, between 3 meters and 5 meters, etc.), the MR system may use a variable depth grid to perform re-projection. On the other hand, if the user is viewing a far-distance scene that exceeds a certain threshold (e.g., farther than 5 meters, 7 meters, or 10 meters), the scene is effectively at infinity, and therefore, the MR system may choose to use a planar constant depth grid instead.
[0098] exist Figures 5 to 7 Various embodiments of the methods outlined in
[0045] describe displaying further insights into a scene to a user. These methods may be implemented by Figure 1 1 and 2. The system 100 shown in FIG. 1 may be implemented with a system similar or identical to the system 100 shown in FIG.
[0099] Figure 5 An embodiment of a method 500 designed to generate a perspective scene and present the perspective scene to a user is shown. In step 510, the method 500 includes receiving image data of a scene for display to the user; the image data includes a left image for the left eye and a right image for the right eye. As previously described, the left image can be a left camera (e.g., Figure 2A The left camera 220L shown in FIG. 220B may capture an image of a scene, while the right image may be captured by a right camera (e.g., Figure 2A In some instances, these left and right images may constitute image frames of a video feed that is systematically collected at a selected frame rate (e.g., 24, 30, 60, 90, or 120 frames per second). In some cases, a lighting module (e.g., Figure 1 The light emitting module 160 in the image processing module is used to improve the lighting of the environment to improve image acquisition.
[0100] Furthermore, at step 515, method 500 includes determining the user's gaze direction and / or eye convergence via an eye tracking module. The eye tracking module may include a camera integrated into the user's head-mounted display (HMD) and directed toward the user's eyes. The eye tracking module may calculate the user's gaze direction and / or eye convergence by analyzing reflections from the user's eyes captured by the camera.
[0101] At step 520, method 500 includes using gaze direction and / or eye convergence to identify an object in the scene that the user is focusing on. This can be done by identifying the line of sight (e.g., left line of sight 306L and right line of sight 306R, as shown in FIG. Figure 3B ) intersect in space to identify the object, thereby identifying the object that the user is focusing on. In some embodiments, object identification may be ambiguous because the system may infer that the object of interest may be located at the user's gaze convergence location. At step 525, method 500 includes determining the left depth from the left eye to the object and the right depth from the right eye to the object. The left depth and right depth can be identified by a triangulation procedure using the angles after the left eye and the right eye are rotated relative to the normal direction. For example, Figure 3B As shown, the angle θ can be used L and θ R , and distance D LR, determine the sides of the triangle formed between the object 318, the gaze direction of the user's left eye 305L, and the gaze direction of the right eye 305R, thereby determining the left depth representing the distance from the left eye 305L to the object 318 and the right depth representing the distance from the right eye 305R to the object 318. In some cases, in addition to or instead of the triangulation process, a suitable depth sensor can be used to determine the left depth and the right depth. A suitable depth sensor may include a time-of-flight depth sensor or any similar sensor as described in detail in the system 100 previously described. In one embodiment, the MR system can project the user's eye gaze into the scene and calculate the intersection between the scene and the user's left gaze and the intersection between the scene and the user's right gaze based on the scene depth information collected using the depth sensor. The distance between each intersection point and each of the user's eyes can then be calculated and used as the left depth and right depth of the user's eyes. For example, the depth sensor can be configured to use the depth sensor along, for example, Figure 3B The depth value of the object is determined by the gaze direction 308 shown. The determined depth value is then forwarded to a computing device, such as Figure 1 The computing device 101 shown in FIG. The computing device can be configured to calculate the left depth to the object, including the depth value, the gaze direction, and / or the vector between the left eye and the depth sensor. In addition, the computing device is configured to determine the right depth to the object by utilizing the depth value, the gaze direction, and / or the vector between the right eye and the depth sensor.
[0102] At step 530, method 500 includes generating a left depth mesh having a constant depth for the left eye based on the left depth, and at step 535, generating a right depth mesh having a constant depth for the right eye based on the right depth. The left depth mesh and the right depth mesh may constitute segments of a spherical mesh, each segment maintaining a consistent left depth and right depth. One of ordinary skill in the art will recognize that the order in which the two depth meshes are generated may vary (e.g., the two depth meshes may be generated in parallel or in any order).
[0103] At step 540, method 500 includes generating a left output image by projecting the left image onto a left depth grid and projecting it to the user's left eye, and generating a right output image by projecting the right image onto a right depth grid and projecting it to the user's right eye at step 545. Those skilled in the art will recognize that the order in which the two output images are generated can vary (e.g., the two output images can be generated in parallel or in any order). Projecting the image captured by the camera onto the depth grid can involve mapping pixel information from the image to corresponding locations in the depth grid. This mapping is typically performed based on a spatial relationship established between pixels in the image and corresponding points or vertices in the depth grid. Projecting the image onto the depth grid can involve aligning pixel coordinates in the image with spatial coordinates in the depth grid. The image captured by the camera is effectively used as a texture for the depth grid. The projected left and right images are then rendered from the viewpoints of the user's left and right eyes, respectively (i.e., the captured images are re-projected to the user's eyes).
[0104] At step 550, the method 500 displays a left output image for the user's left eye. At step 555, the method displays a right output image for the user's right eye. Figure 1 The display module 150 of the illustrated system 100 is used by a suitable display module to render the displayed image.
[0105] Figure 6 Another embodiment of a method 600 for generating and displaying a see-through scene to a user is shown. Method 600 employs a hybrid approach that uses a constant depth grid to reproject a portion of an acquired image and a variable depth grid to reproject another portion of the image. Method 600 includes steps 610 through 635, which may be similar or identical to steps 510 through 535 of method 500. Additionally, method 600 includes a step 637 of determining a variable depth grid for the scene using a depth estimation module. For example, the variable depth grid for the scene may be the same as Figure 3A The depicted variable depth grid 343 is very similar or identical.In some scenarios, the variable depth grid 343 may be acquired by utilizing a depth sensor that is configured to scan the user's environment for the purpose of determining the depth grid.
[0106] At step 639, the method 600 generates a first left output image by reprojecting a first portion of the left image using the left constant depth grid. The first portion of the left image may be a particular portion of the left image that includes details very close to the object that the user is focusing on. By way of illustration, the first portion of the left image may include a defined viewing cone (e.g., Figure 3D Image data of all objects within the viewing cone 333).
[0107] Similar to step 639, at step 641, method 600 generates a first right output image by reprojecting a first portion of the right image using the right constant depth grid. Similar to step 639, the designated portion of the right image encapsulates details near the object the user is focusing on. The spatial extent is typically defined by a related viewing cone similar to that mentioned in step 639.
[0108] At step 643, method 600 generates a second left output image by reprojecting a second portion of the left image using the variable depth grid, the second portion of the left image being complementary to the first portion of the left image. For example, the second portion of the left image may include a portion located at a position such as Figure 3D The image data for all objects outside of the viewing cone 333 is shown. In some cases, the second portion of the left output image may be blurred before re-projection to indicate that the objects represented by the image data are further away from the user.
[0109] Similar to step 643, at step 645, method 600 generates a second right output image by reprojecting a second portion of the right image using the variable depth grid, the second portion of the right image being complementary to the first portion of the right image. Similar to step 643, the second designated portion of the right image encapsulates the image data such as Figure 3D The image data for all objects outside the viewing cone 333 is shown. In some cases, the second portion of the right output image may be blurred before re-projection to indicate that the objects represented by the image data are further away from the user.
[0110] In some cases, if the first portion of the left (or right) image includes image data corresponding to an object within the first viewing frustum, the second portion of the left (or right) image may include image data for an object located outside the second viewing frustum but within the boundaries of the first viewing frustum. This selection of the first portion and the second portion facilitates an intentional overlap between the first portion of the left (or right) image and the second portion of the left (or right) image, thereby ensuring a cohesive transition between the first portion and the second portion when the projections of these portions are subsequently combined in step 647 (or step 649) of method 600, as will be further described below.
[0111] At steps 647 and 649, method 600: (1) combines the generated first left output image and the second left output image to generate a combined left output image (step 647) and (2) combines the generated first right output image and the second right output image to generate a combined right output image (step 649). As described above, in some cases, the first and second portions of the left image may overlap in at least some areas to improve the subsequent combination of the generated first and second left output images into the combined left output image. Similarly, the first and second portions of the right image may overlap in at least some areas to improve the subsequent combination of the generated first and second right output images into the combined right output image.
[0112] When different parts of the images are projected onto one or more depth grids, the process of combining or stitching these images can involve aligning and blending overlapping or touching areas to create a seamless and coherent final image. Such an image stitching process can include feature matching, which involves identifying significant features in the overlapping areas of adjacent images. These features can include key points, corners, or edges. In addition, the stitching process can include image alignment and / or applying geometric transformations (e.g., translation, rotation, or scaling) to properly align the images. Transformation matrices can be used for this purpose. In addition, stitching can include blending overlapping areas to eliminate visible seams, for example by adjusting pixel intensities at boundaries to create a seamless transition between adjacent images. In addition, stitching can include color correction by ensuring consistency of color and brightness across the stitched images. In addition, color correction techniques can be applied to improve the overall appearance. Similarly, color correction can be used to ensure that the left output image matches the right output image in color, and / or brightness, and / or sharpness.
[0113] At steps 651 and 653, method 600 displays (1) a combined left output image for the user's left eye and (2) a combined right output image for the user's right eye, respectively. Figure 1 The display module 150 of the illustrated system 100 is used by a suitable display module to render the displayed image.
[0114] Figure 7A method 700 for adaptively selecting different types of depth grids for perspective generation is shown. Steps 710 to 725 may be similar or identical to corresponding steps 510 to 525 of method 500 or corresponding steps 610 to 625 of method 600. Furthermore, at step 727, method 700 may evaluate whether at least one of the left depth or the right depth is below a depth threshold (i.e., the depth is sufficiently close). The depth threshold may be selected in various ways and may have any suitable value, such as tens of centimeters, half a meter, one meter, several meters, or any distance ranging from a few centimeters to several meters. In some instances, the depth threshold may be adjusted for each user's specific circumstances. Alternatively, the depth threshold may be determined based on an overall reduction in artifacts when presenting a scene to the user. When either the left depth or the right depth is below the depth threshold (step 727, yes), method 700 may determine that the constant depth grid is suitable for perspective generation. Therefore, method 700 may proceed to steps 730 to 755, which may be similar or identical to corresponding steps 530 to 555 of method 500. On the other hand, when neither the left depth nor the right depth is below the depth threshold (step 727, no), method 700 may determine that a variable depth grid is more appropriate for the scene. Therefore, method 700 may proceed to step 732, where a variable depth grid for the scene is determined using a depth estimation module. Step 732 may be similar or identical to step 637 of method 600. Furthermore, after step 732 is completed, method 700 includes generating a left output image by reprojecting the left image using the variable depth grid at step 734, and generating a right output image by reprojecting the right image using the variable depth grid at step 736. Subsequently, method 700 may proceed to steps 750 and 755, which may be similar or identical to steps 550 and 555 of method 500.
[0115] The scope of the present disclosure covers all changes, substitutions, variations, alterations and modifications to the example embodiments described or shown herein, which will be understood by those skilled in the art. The scope of the present disclosure is not limited to the example embodiments described or shown herein. In addition, although the present disclosure describes and illustrates the various embodiments herein as including specific parts, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any of the parts, elements, features, functions, operations, or steps described or shown anywhere herein as will be understood by those skilled in the art. In addition, in the appended claims, references to a device or system that is adapted to, arranged to, capable of, configured to, enabled to, operable to, or operable to perform a particular function, or a component in a device or system, include the device, system, component, regardless of whether the device, system, component, or the particular function is activated, turned on, or unlocked, as long as the device, system, or component is so adapted to, arranged to, capable of, configured to, enabled to, operable to, or operable. Additionally, although this disclosure may describe or illustrate a particular embodiment as providing particular advantages, that particular embodiment may not provide these advantages, or may provide some or all of these advantages.
[0116] As used herein, unless expressly stated otherwise or the context indicates otherwise, "or" is inclusive, not exclusive. Thus, as used herein, "A or B" means "A, B, or both," unless expressly stated otherwise or the context indicates otherwise. Furthermore, "and" is both joint and several. Thus, as used herein, "A and B" means "A and B, jointly or severally," unless expressly stated otherwise or the context indicates otherwise.
Claims
1. A method for displaying a scene to a user, the method comprising, by a computing device: receiving image data of the scene for display to the user, the image data comprising a left image for a left eye and a right image for a right eye; determining the user's gaze direction or eye convergence through an eye tracking module; identifying an object in the scene that the user is focusing on using the gaze direction or the eye convergence; determining a left depth from the left eye to the object and a right depth from the right eye to the object; generating a left depth grid having a constant depth for the left eye based on the left depth; generating a right depth grid having a constant depth for the right eye based on the right depth; generating a left output image by projecting the left image onto the left depth grid; generating a right output image by projecting the right image onto the right depth grid; displaying the left output image for the left eye; as well as The right output image is displayed for the right eye.
2. The method according to claim 1, wherein The computing device is communicatively connected to a left camera and a right camera of a head-mounted device worn by the user, and wherein the left image is acquired by the left camera and the right image is acquired by the right camera.
3. The method according to claim 1, wherein The eye tracking module includes cameras directed toward the left eye and the right eye of the user.
4. The method according to claim 1, wherein The position of the object is determined based on the eye convergence of the user.
5. The method according to claim 4, wherein The position of the object is also determined based on scene information.
6. The method according to claim 1, wherein Identifying the object in the scene includes determining an intersection between the gaze direction of the user and a depth of the scene.
7. The method according to claim 1, wherein Identifying the object in the scene includes determining an intersection between the gaze direction of the user and a 3D model of the scene.
8. The method according to claim 1, wherein Identifying the object in the scene is also based on a current usage context of the computing device.
9. The method according to claim 1, wherein The left depth grid and the right depth grid are spherical.
10. The method according to claim 1, wherein The left depth grid and the right depth grid are planar.
11. The method according to claim 1, wherein The constant depth of the left depth grid is different from the constant depth of the right depth grid.
12. The method according to claim 1, further comprising: After generating the left depth grid and the right depth grid, identifying a second object in the scene that the user is focusing on using a second gaze direction or a second eye convergence of the user, wherein the second object is different from the object; determining a second left depth from the left eye to the second object and a second right depth from the right eye to the second object; generating a second left depth grid having a constant depth for the left eye based on the second left depth; generating a second right depth grid having a constant depth for the right eye based on the second right depth; generating a second left output image using the second left depth grid; generating a second right output image using the second right depth grid; displaying the second left output image for the left eye; and The second right output image is displayed for the right eye.
13. One or more computer-readable non-transitory storage media containing software that, when executed, is operable to: receiving image data of a scene surrounding a user, the image data comprising a left image for a left eye and a right image for a right eye; determining the user's gaze direction or eye convergence through an eye tracking module; identifying a region of interest in the scene using the gaze direction or the eye vergence; determining a left depth from the left eye to the region of interest and a right depth from the right eye to the region of interest; generating a left depth grid having a constant depth for the left eye based on the left depth; generating a right depth grid having a constant depth for the right eye based on the right depth; generating a left output image by projecting the left image onto the left depth grid; generating a right output image by projecting the right image onto the right depth grid; displaying the left output image for the left eye; as well as The right output image is displayed for the right eye.
14. The one or more computer-readable non-transitory storage media of claim 13, wherein: The location of the region of interest is determined based on the eye convergence of the user.
15. The one or more computer-readable non-transitory storage media of claim 14, wherein: The position of the region of interest is also determined based on scene information.
16. The one or more computer-readable non-transitory storage media of claim 13, wherein: The left depth grid and the right depth grid are spherical.
17. A system comprising: one or more processors; as well as one or more computer-readable non-transitory storage media coupled to one or more of the one or more processors and storing instructions that, when executed by one or more of the one or more processors, are operable to cause the system to: receiving image data of a scene surrounding a user, the image data comprising a left image for a left eye and a right image for a right eye; determining the user's gaze direction or eye convergence through an eye tracking module; identifying a region of interest in the scene using the gaze direction or the eye vergence; determining a left depth from the left eye to the region of interest and a right depth from the right eye to the region of interest; generating a left depth grid having a constant depth for the left eye based on the left depth; generating a right depth grid having a constant depth for the right eye based on the right depth; generating a left output image by projecting the left image onto the left depth grid; generating a right output image by projecting the right image onto the right depth grid; displaying the left output image for the left eye; as well as The right output image is displayed for the right eye.
18. The system according to claim 17, wherein: The location of the region of interest is determined based on the eye convergence of the user.
19. The system according to claim 18, wherein: The position of the region of interest is also determined based on scene information.
20. The system of claim 17, wherein: The left depth grid and the right depth grid are spherical.