Virtual, Augmented, and Mixed Reality Systems and Methods
By tracking the user's line of sight and analyzing 3D data to identify virtual objects, the VR/AR/MR system accurately determines the user's depth of focus, addressing the vergence-accommodation conflict and enhancing the 3D viewing experience.
Patent Information
- Application Number
- JP2022533463
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-06
- Filing Date
- 2020-12-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-12-03
AI Technical Summary
Current VR/AR/MR systems face challenges in accurately determining the depth of focus for users, leading to sub-optimal 3D viewing experiences due to the vergence-accommodation conflict and limitations in eye-tracking hardware and processing power.
The method involves tracking the user's line of sight and analyzing 3D data to identify virtual objects along this line. By determining the depth of the virtual object that intersects the user's line of sight, the system can accurately identify the user's depth of focus, which can be further refined through convergence/divergence motion analysis.
This approach enhances the accuracy of determining the user's depth of focus, reducing the vergence-accommodation conflict and improving the overall 3D viewing experience while efficiently allocating limited hardware resources.
Smart Images

Figure 0007693673000001 
Figure 0007693673000002 
Figure 0007693673000003
Abstract
Description
Technical Field
[0001] (Copyright Notice) Part of the disclosure of this patent document contains materials that should be protected by copyright. The copyright owner has no objection to anyone copying this patent document or this patent disclosure as long as it appears in the patent file or records of the Patent and Trademark Office, but in other cases, the copyright owner retains all copyrights.
[0002] This disclosure relates to virtual reality, augmented reality, and mixed reality imaging, visualization, and display systems and methods. In particular, this disclosure relates to virtual reality, augmented reality, and mixed reality imaging, visualization, and display systems and methods for determining the depth of focus of a viewer.
Background Art
[0003] Modern computing and display technologies are facilitating the development of virtual reality (VR), augmented reality (AR), and mixed reality (MR) systems. VR systems create a simulated environment for a user to experience. This can be done by presenting computer-generated images to the user through a head-mounted display. This image creates a sensory experience that immerses the user in the simulated environment. VR scenarios typically involve not only the presentation of computer-generated images but also real-world images.
[0004] AR systems generally complement the real-world environment with simulated elements. For example, an AR system may provide a user with a view of the surrounding real-world environment via a head-mounted display. However, computer-generated images can also be presented on the display and can enhance the real-world environment. The computer-generated images can include elements that are contextually relevant to the real-world environment. Such elements can include simulated text, images, objects, etc. MR systems also introduce simulated objects into the real-world environment, but these objects typically feature a higher degree of interactivity than those in AR systems. The simulated elements can often be two-way in real time.
[0005] Figure 1 depicts an exemplary AR / MR scene 2, where the user sees a real-world park setting 6 featuring people, trees, buildings in the background, and a concrete platform 20. In addition to these items, computer-generated images are also presented to the user. The computer-generated images can include, for example, a robot image 10 standing on the real-world platform 20 and an avatar character 12 in the form of a flying cartoon that appears like an anthropomorphic bumblebee, but these elements 12, 10 do not actually exist within the real-world environment.
[0006] Various optical systems generate images at various depths to display VR, AR, or MR scenarios.
[0007] A VR / AR / MR system can operate in various display modes, including a blended light field and a discrete light field mode. In the blended light field mode (the "blend mode"), the display presents a composite light field that includes images displayed on two or more discrete depth planes simultaneously or nearly simultaneously. In the discrete light field mode (the "discrete mode"), the display presents a composite light field that includes an image displayed on a single depth plane. The illuminated depth plane is controllable in the discrete mode. Some VR / AR / MR systems include an eye tracking subsystem for determining the viewer's focus position. Such systems can illuminate the depth plane corresponding to the viewer's focus position.
[0008] Because the human visual perception system is complex, it is difficult to produce VR / AR / MR technologies that facilitate a comfortable, natural, and rich presentation of virtual image elements among other virtual or real-world image elements. Three-dimensional ("3D") image display systems suffer from the vergence-accommodation conflict problem. This problem occurs when two biologically related processes related to optical depth send conflicting depth signals to the user / viewer's brain. Vergence is related to the viewer's eye tendency to rotate to align the object of the viewer's attention at a certain distance with the optical axis (axis). In a binocular system, the point where the optical axes intersect can be called the "vergence point". The amount of rotation of the viewer's eyes during vergence is interpreted by the viewer's brain as the estimated depth. Accommodation is related to the viewer's eye lens tendency to focus on the object of the viewer's attention at a certain distance. The focus of the viewer's eyes during vergence is interpreted by the viewer's brain as another estimated depth. When the vergence and accommodation signals are interpreted by the viewer's brain as the same or similar estimated depth, the 3D viewing experience is natural and comfortable for the viewer. On the other hand, when the vergence and accommodation signals are interpreted by the viewer's brain as substantially different estimated depths, the 3D viewing experience is sub-optimal for the viewer and can cause discomfort (such as eye strain, headache, etc.) and fatigue. Such a problem is known as the vergence-accommodation conflict.
[0009] Portable VR / AR / MR systems also have other limitations such as size and portability issues, battery life issues, system overheating issues, processing power, memory, bandwidth, data sources, component latency, and other system and optical challenges, which can negatively impact VR / AR / MR system performance. These other limitations must be considered when implementing 3D image rendering techniques for natural convergence / divergence motion and depth adjustment. To accommodate these hardware-related limitations of VR / AR / MR systems, it is often useful to determine the user / viewer's focus. Some systems map the user / viewer's focus to the user / viewer's convergence / divergence motion points, but limitations in eye-tracking hardware and high-speed user / viewer eye movements reduce the accuracy of determining the user / viewer's focus using this method. The step of determining the user / viewer's focus enables the VR / AR / MR system to effectively allocate limited hardware resources. In particular, determining the depth of the user / viewer's focus enables the VR / AR / MR system to direct limited hardware resources to a specific plane that coincides with the user / viewer's focus.
[0010] One way to determine the depth of focus of a user / viewer is to use depth perception to generate a depth map of the entire field of view and then use eye tracking to identify the depth of focus of the user / viewer using the depth map of the field of view. Depth perception is the determination of the distance between known points in a three-dimensional ("3D") space (e.g., a sensor) and points of interest ("POIs") on the surface of an object. Depth perception also includes the step of determining the individual distances of multiple POIs on the surface, which is also known as texture perception because it determines the texture of the surface. Depth or texture perception is useful for many computer vision systems, including augmented reality systems. The step of generating a depth map can be a demanding process in terms of hardware (e.g., infrared cameras and infrared light projectors) and processing power. In a mobile VR / AR / MR system, the depth map will need to be continuously updated at some cost in terms of processing power, memory, communication channels, etc.
[0011] For example, systems and techniques for determining the depth of focus of a user / viewer for rendering and displaying a 3D image to the viewer / user while minimizing vergence-accommodation conflict, and systems and techniques for minimizing the requirements on the limited graphical processing capabilities of a portable VR / AR / MR system while doing so, are needed. Improved systems and techniques for processing image data and displaying images are needed to address these issues. The systems and methods described herein are configured to address these and other issues.
[0012] What is needed are techniques or a plurality of techniques for improving conventional techniques and / or other approaches under consideration. Some of the approaches described in this background section are approaches that could be pursued, but are not necessarily approaches that have been previously contemplated or pursued. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0013] In one embodiment, a method for determining the depth of focus of a user of a three-dimensional ("3D") display device includes tracking a first line of sight of the user. The method also includes analyzing 3D data to identify one or more virtual objects along the user's first line of sight. The method further includes identifying the depth of the only virtual object that intersects the user's first line of sight as the user's depth of focus when only one virtual object intersects the user's first line of sight.
[0014] In one or more embodiments, each of one or more virtual objects is displayed at an individual depth from the user. The step of analyzing the plurality of 3D data may include analyzing depth segmentation data corresponding to the 3D data. The method may also include performing a convergence / divergence motion analysis on the user to generate a convergence / divergence motion depth of focus of the user and using the convergence / divergence motion depth of focus of the user to improve the accuracy of the user's depth of focus.
[0015] In one or more embodiments, when more than one virtual object intersects the user's first line of sight, the method may also include tracking a second line of sight of the user. The method may also include analyzing 3D data to identify one or more virtual objects along the user's second line of sight. The method further may include identifying the depth of the only 2D image frame as the user's depth of focus when only one virtual object intersects both the user's first line of sight and the second line of sight.
[0016] In another embodiment, a method for determining the depth of focus of a user of a three-dimensional ("3D") display device includes tracking the user's first line of sight. The method also includes analyzing a plurality of two-dimensional ("2D") image frames displayed by the display device to identify one or more virtual objects along the user's first line of sight. The method further includes identifying the depth of only one of the 2D image frames as the user's depth of focus when only one of the plurality of 2D image frames includes a virtual object along the user's first line of sight.
[0017] In one or more embodiments, each of the plurality of 2D image frames is displayed at an individual depth from the user. Analyzing the plurality of 2D image frames may include analyzing depth segmentation data corresponding to the plurality of 2D image frames. The method may also include performing a convergence / divergence motion analysis on the user to generate a convergence / divergence motion depth of focus of the user, and using the convergence / divergence motion depth of focus of the user to improve the accuracy of the user's depth of focus.
[0018] In one or more embodiments, when more than one of the plurality of 2D image frames includes a virtual object along the user's first line of sight, the method may also include tracking the user's second line of sight. The method may also include analyzing a plurality of 2D image frames displayed by the display device to identify one or more virtual objects along the user's second line of sight. The method further includes identifying the depth of only one of the 2D image frames as the user's depth of focus when only one of the plurality of 2D image frames includes a virtual object along both the user's first line of sight and the user's second line of sight. The present invention provides, for example, the following. (Item 1) A method for determining the depth of focus of a user of a three-dimensional (``3D'') display device, tracking a first line of sight path of the user, analyzing 3D data to identify one or more virtual objects along the first line of sight path of the user, identifying the depth of the only one virtual object as the depth of focus of the user when only one virtual object intersects the first line of sight path of the user A method comprising: (Item 2) The method according to item 1, wherein each of the one or more virtual objects is displayed at an individual depth from the user. (Item 3) The method according to item 1, wherein analyzing the plurality of 3D data includes analyzing depth segmentation data corresponding to the 3D data. (Item 4) When more than one virtual object intersects the first line of sight path of the user, tracking a second line of sight path of the user, analyzing 3D data to identify the one or more virtual objects along the second line of sight path of the user, identifying the depth of only one 2D image frame as the depth of focus of the user when only one virtual object intersects both the first and second lines of sight paths of the user The method according to item 1, further comprising: (Item 5) performing a convergence / divergence motion analysis on the user to generate a convergence / divergence motion depth of focus of the user, using the convergence / divergence motion depth of focus of the user to improve the accuracy of the depth of focus of the user The method according to item 1, further comprising: (Item 6) A method for determining the depth of focus of a user of a three-dimensional (``3D'') display device, tracking a first line of sight path of the user, analyzing a plurality of two-dimensional (``2D'') image frames displayed by the display device to identify one or more virtual objects along the first line of sight path of the user, identifying the depth of only one of the plurality of 2D image frames as the depth of focus of the user when only one of the plurality of 2D image frames includes a virtual object along the first line of sight path of the user A method comprising: (Item 7) The method according to item 6, wherein each of the plurality of 2D image frames is displayed at an individual depth from the user. (Item 8) The method according to item 6, wherein analyzing the plurality of 2D image frames includes analyzing depth segmentation data corresponding to the plurality of 2D image frames. (Item 9) When more than one of the plurality of 2D image frames includes a virtual object along a first line of sight path of the user, tracking a second line of sight path of the user; and analyzing the plurality of 2D image frames displayed by the display device to identify one or more virtual objects along the second line of sight path of the user; identifying the depth of only the one 2D image frame as the focus depth of the user when only one of the plurality of 2D image frames includes a virtual object along both the first and second lines of sight paths of the user. The method according to item 6, further comprising. (Item 10) performing a convergence / divergence motion analysis on the user to generate a convergence / divergence motion focus depth of the user; and using the convergence / divergence motion focus depth of the user to improve the accuracy of the focus depth of the user. The method according to item 6, further comprising.
Brief Description of the Drawings
[0019] This patent or patent application contains at least one drawing created in color. Copies of this patent or patent application publication with color drawings will be provided by the Patent Office upon request and payment of the required fees.
[0020] The drawings described below are for illustrative purposes only. The drawings are not intended to limit the scope of the present disclosure. This patent or patent application contains at least one drawing created in color. Copies of this patent or patent application publication with color drawings will be provided by the United States Patent and Trademark Office upon request and payment of the required fees.
[0021] The drawings illustrate the design and utility of various embodiments of the present disclosure. Note that the figures are not drawn to exact scale and elements of similar structure or function are represented by like reference numerals throughout the figures. A more detailed description of the present disclosure will be provided by reference to the specific embodiments illustrated in the accompanying drawings in order to gain a deeper understanding of the various listed and other advantages and objects of the present disclosure. It is understood that these drawings depict only typical embodiments of the present disclosure and should not be considered as limiting its scope, and that the present disclosure will be described and explained with additional specificity and detail through the use of the accompanying drawings.
[0022]
Figure 1
[0023]
Figure 2
Figure 3
Figure 4
Figure 5
[0024]
Figure 6
[0025]
Figure 7
[0026]
Figure 8
[0027]
Figure 9
[0028]
Figure 10A
[0029]
Figure 10B
[0030]
Figure 11
[0031]
Figure 12A
[0032]
Figure 12B
[0033]
Figure 13
Figure 14
[0034]
Figure 15
[0035]
Figure 16
[0036]
Figure 17
[0037] Detailed Description Various embodiments of the present disclosure are directed to systems, methods, and articles of manufacture for virtual reality (VR) / augmented reality (AR) / mixed reality (MR) in a single embodiment or multiple embodiments. Other objects, features, and advantages of the present disclosure are set forth in the detailed description, figures, and claims.
[0038] Various embodiments will now be described in detail with reference to the drawings, which are provided as illustrative examples to enable those skilled in the art to practice the present disclosure. It should be noted that the figures and the following examples are not meant to limit the scope of the present disclosure. Some elements of the present disclosure may be implemented, in part or in whole, using known components (or methods or processes), and only those portions of such known components (or methods or processes) necessary for an understanding of the present disclosure will be described, with detailed descriptions of other portions of such known components (or methods or processes) being omitted so as not to obscure the present disclosure. Further, various embodiments include presently known and future known equivalents of components referred to herein by way of example.
[0039] Embodiments according to the present disclosure often address implementation issues of VR / AR / MR systems that rely on a combination of off-the-shelf components and custom components. In some cases, the off-the-shelf components do not possess all of the features or performance characteristics required to implement a desired aspect of the VR / AR / MR system to be deployed. Some embodiments target an approach for adding the ability to adapt to desired features or performance characteristics of the VR / AR / MR system to be deployed and / or for reusing resources for another purpose. The accompanying figures and discussion herein present an exemplary environment, system, method, and computer program product for VR / AR / MR systems.
[0040] The head-mounted visual display system and the 3D image rendering system may be implemented independently of the AR / MR system, but some of the following embodiments are described in the context of the AR / MR system for illustrative purposes only. The 3D image rendering and display systems described herein may also be used in a similar manner for VR systems. Overview of the Problem and Solution
[0041] As described above, VR / AR / MR systems also have limitations such as size and portability, battery life, system overheating, processing power, memory, bandwidth, data sources, component latency, and other system and optical issues, which can negatively impact VR / AR / MR system performance. These limitations result in reduced graphical processing and image display requirements and can pose countervailing challenges to improving 3D image rendering and display. VR / AR / MR systems can more effectively deploy limited processing and display resources by identifying the depth plane corresponding to the focus of the user / viewer. However, using only convergence / divergence motion to identify the focus plane of the user / viewer may not be accurate enough. On the other hand, the step of generating a depth map for use in identifying the focus plane of the user / viewer can add processing and hardware requirements to the VR / AR / MR system.
[0042] For example, due to potential graphical processing and size / portability issues, VR / AR / MR systems, particularly head-mounted systems, may include only sufficient components to render and display a color image on one depth plane per frame at a minimum frame rate (e.g., 60 Hz) for smooth display of moving virtual objects (i.e., discrete mode). An example of such a VR / AR / MR system operating in discrete mode is shown in FIGS. 8A-9B. As schematically shown in FIG. 8A, a 3D image includes three virtual objects (near and far cubes 810, 812 and one cylinder 814) adjacent to various depth planes (near cube 810 adjacent to the vicinity of depth plane 0 and far cube 812 and cylinder 814 adjacent to far depth plane 1). In some embodiments, the near depth plane 0 is at about 1.96 diopters and the far depth plane 1 is at about 0.67 diopters. FIG. 8B is a perspective view of the viewer of the 3D image shown in FIG. 8A. In FIG. 8B, eye tracking of user / viewer 816 shows that the viewer's eyes converge to the convergence / divergence movement point 818 that coincides with the location of the near cube 810. In discrete mode, only a single depth plane is illuminated per frame (i.e., the image is rendered and displayed). As schematically shown in FIG. 9A, since the convergence / divergence movement point 818 coincides with the location of the near cube 810 adjacent to the vicinity of depth plane 0, all the content of the 3D image (i.e., near and far cubes 810, 812 and cylinder 814) is projected onto the near depth plane 0. FIG. 9B is a perspective view of the viewer of the 3D image after its content has been projected onto the near depth plane 0. Only the near depth plane 0 is illuminated and the eyes of viewer 816 accommodate to the near depth plane 0. As explained above, the equalization of the user / viewer's focus and the user / viewer's convergence / divergence movement point may result in a less accurate determination of this critical data point.
[0043] In some embodiments, all projections of the 3D image content onto a single depth plane induce only minimal vergence-accommodation conflict (e.g., minimal user discomfort, eye strain, headache). This is because there is a loose coupling between accommodation and vergence, such that the human brain will tolerate a mismatch of up to approximately 0.75 diopters between accommodation and vergence. As shown in FIG. 10A, this ±0.75 diopter tolerance translates into near-accommodation zone 1010 and far-accommodation zone 1012. Due to the inverse relationship between distance and diopters, as shown in FIG. 10B, far-accommodation zone 1012 is larger than the nearer near-accommodation zone 1010. In the embodiment depicted in FIG. 10A, by using near depth plane 0 and far depth plane 1, the ±0.75 diopter tolerance also results in a vergence-accommodation zone overlap 1014 where object depths corresponding within the vergence-accommodation zone overlap 1014 can be displayed on one or both of near depth plane 0 and far depth plane 1, with different brightness and / or color values, for example, at different scales. In embodiments where all of the 3D image content is located within either near-accommodation zone 1010 or far-accommodation zone 1012 and the viewer 816's eyes converge on that depth plane, all projections of the 3D image content onto that depth plane induce only minimal vergence-accommodation conflict.
[0044] Operation in discrete mode enables the VR / AR / MR system to deliver 3D images while functioning within its hardware limitations (e.g., processing, display, memory, communication channels, etc.). However, a discrete mode display requires accurate identification of the single depth plane to be illuminated while minimizing demands on system resources. VR / AR / MR systems operating in blend mode can also reduce demands on system resources by focusing their resources on the depth plane corresponding to the user / viewer's focus. Again, this resource-saving method for blend mode displays requires accurate identification of the user / viewer's focal plane.
[0045] The embodiments described herein include 3D image rendering and display systems and methods for use in conjunction with various VR / AR / MR systems. These 3D image rendering and display systems and methods identify depth planes corresponding to the focus of the user / viewer for use in both discrete mode and blended mode displays while reducing the system resources consumed, thereby addressing many of the problems described above. Exemplary VR, AR, and / or MR systems
[0046] The following description relates to exemplary VR, AR, and / or MR systems in which embodiments of various 3D image rendering and display systems may be practiced using the same. However, it should be understood that the embodiments are also suitable for use in other types of display systems (including other types of VR, AR, and / or MR systems), and thus the embodiments are not limited to only the exemplary systems disclosed herein.
[0047] The VR / AR / MR system disclosed in this specification can include a display that presents computer-generated images (video / image data) to the user. In some embodiments, the display system is wearable, which can advantageously provide a more immersive VR / AR / MR experience. Various components of the VR, AR, and / or MR virtual image system 100 are depicted in FIGS. 2-5. The virtual image generation system 100 includes a frame structure 102 worn by an end user 50, a display subsystem 110 carried by the frame structure 102 such that the display subsystem 110 is positioned in front of the eyes of the end user 50, and a speaker 106 carried by the frame structure 102 such that the speaker 106 is positioned adjacent to the outer ear canal of the end user 50 (optionally, another speaker (not shown) is positioned adjacent to the other outer ear canal of the end user 50 to provide stereo / formable sound control). The display subsystem 110 is designed to present a light pattern that can be comfortably perceived by the eyes of the end user 50 as an augmentation to the physical reality with a high level of image quality and three-dimensional perception and can present two-dimensional content. The display subsystem 110 presents a sequence of frames at a high frequency to provide the perception of a single coherent scene.
[0048] In the illustrated embodiment, the display subsystem 110 employs an "optical see-through" display through which a user can directly view light from real objects through a transparent (or translucent) element. The transparent element is often referred to as a "combiner" and superimposes light from the display across the field of view of a real-world user. To achieve this purpose, the display subsystem 110 includes a partially transparent display. In some embodiments, the transparent display may be electronically controlled. In some embodiments, the transparent display includes segmented dimming and may control the transparency of one or more portions of the transparent display. In some embodiments, the transparent display includes global dimming and may control the overall transparency of the transparent display. The display is positioned in the field of view of the end user 50 between the end user 50's eye and the surrounding environment such that direct light from the surrounding environment is transmitted through the display to the end user 50's eye.
[0049] In the illustrated embodiment, the image projection assembly provides light to a partially transparent display, which is thereby combined with direct light from the ambient environment and transmitted from the display to the eyes of user 50. The projection subsystem may be a fiber optic scanning-based projection device, and the display may be a waveguide-based display into which scanned light from the projection subsystem is input, for example, at a single optical viewing distance that is closer to infinity (e.g., arm's length), an image at a plurality of discrete optical viewing distances or focal planes, and / or an image layer that is stacked at a plurality of viewing distances or focal planes and represents a three-dimensional 3D object. These layers in the light field may be stacked sufficiently close to each other to appear contiguous to the human visual system (i.e., one layer is within the cone of confusion of an adjacent layer). Additionally or alternatively, photo elements (i.e., sub-images) may be blended across two or more layers, even if those layers are stacked more sparsely, to increase the perceived continuity of the transition between layers in the light field (i.e., one layer is outside the cone of confusion of an adjacent layer). The display subsystem 110 may be for monocular or binocular use. The image projection system may be capable of generating a blend mode, but system limitations (e.g., power, heat, speed, etc.) may limit the image projection system to a discrete mode with a single viewing distance per frame.
[0050] The virtual image generation system 100 may also include one or more sensors (not shown) mounted on the frame structure 102 to detect the position and movement of the head 54 of the end user 50 and / or the position and interpupillary distance of the eyes of the end user 50. Such sensors may include image capture devices (such as cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, and / or gyroscopes. Many of these sensors operate on the assumption that the frame 102 to which they are attached is substantially fixed to the user's head, eyes, and ears. The virtual image generation system 100 may also include one or more cameras directed at the user to track the user's eyes.
[0051] The virtual image generation system 100 may also include a user orientation detection module. The user orientation module may detect the instantaneous position of the head 54 of the end user 50 (e.g., via sensors coupled to the frame 102) and predict the position of the head 54 of the end user 50 based on the position data received from the sensors. Detection of the instantaneous position of the head 54 of the end user 50 facilitates determination of the specific real object that the end user 50 is looking at, thereby providing an indication of the specific virtual object to be generated in relation to that real object and further providing an indication of the position at which the virtual object is to be displayed. The user orientation module may also track the eyes of the end user 50 based on the tracking data received from the sensors.
[0052] The virtual image generation system 100 may also include a control subsystem that can take any of a variety of forms. The control subsystem includes several controllers, for example, one or more microcontrollers, microprocessors or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers, for example, application specific integrated circuits (ASICs), display bridge chips, display controllers, programmable gate arrays (PGAs), for example, field PGAs (FPGAs), and / or programmable logic controllers (PLUs).
[0053] The control subsystem of the virtual image generation system 100 may include a central processing unit (CPU), a graphics processing unit (GPU), one or more frame buffers, and a three-dimensional database for storing three-dimensional data. The CPU may control the overall operation, while the GPU may render frames from the three-dimensional data stored in the three-dimensional database (i.e., convert the three-dimensional scene into a two-dimensional image) and store these frames in the frame buffer. One or more additional integrated circuits may control the reading of frames into and / or the reading from the frame buffer and the operation of the image projection assembly of the display subsystem 110.
[0054] The various processing components of the virtual image generation system 100 may be physically contained within a distributed subsystem. For example, as illustrated in FIGS. 2-5, the virtual image generation system 100 may include a local processing and data module 130 that is operatively coupled to a local display bridge 142, a display subsystem 110, and sensors by means of a wired conductor or wireless connectivity 136, etc. The local processing and data module 130 may be mounted in various configurations, such as fixedly attached to the frame structure 102 (FIG. 2), fixedly attached to a helmet or cap 56 (FIG. 3), removably attached to the body 58 of the end user 50 (FIG. 4), or removably attached to the waist 60 of the end user 50 in a belt-coupled configuration (FIG. 5). The virtual image generation system 100 may also include a remote processing module 132 and a remote data repository 134 that are operatively coupled to the local processing and data module 130 and the local display bridge 142 by means of wired conductors or wireless connectivity 138, 140, etc., and these remote modules 132, 134 are operatively coupled to each other and become available as resources to the local processing and data module 130 and the local display bridge 142.
[0055] The local processing and data module 130 and the local display bridge 142 may each include a power-efficient processor or controller and digital memory such as flash memory, both of which are captured from sensors and / or potentially processed or read out and then acquired and / or processed using the remote processing module 132 and / or the remote data repository 134 for passage to the display subsystem 110, and may be utilized to assist in the processing, caching, and storage of the data. The remote processing module 132 may include one or more relatively powerful processors or controllers configured to analyze and process data and / or image information. The remote data repository 134 may include a relatively large-scale digital data storage facility, which may be available through other networking configurations in an Internet or "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module 130 and the local display bridge 142, enabling complete autonomy from any remote module.
[0056] The couplings 136, 138, 140 between the various components described above may include one or more wired interfaces or ports for providing wired or optical communication, or one or more wireless interfaces or ports via RF, microwave, IR, etc. for providing wireless communication. In some implementations, all communication may be wired, while in other implementations, all communication may be wireless. Still further implementations may have different choices of wired and wireless communication than those illustrated in FIGS. 2-5. Accordingly, a particular choice of wired or wireless communication should not be considered limiting.
[0057] In some embodiments, the user orientation module is contained within the local processing and data module 130 and / or the local display bridge 142, while the CPU and GPU are contained within the remote processing module. In alternative embodiments, the CPU, GPU, or a portion thereof may be contained within the local processing and data module 130 and / or the local display bridge 142. The 3D database can be associated with the remote data repository 134 or can be locally located.
[0058] Some VR, AR, and / or MR systems use multiple volume phase holograms, surface relief holograms, or light guiding optical elements that incorporate depth plane information to generate images that appear to arise from individual depth planes. In other words, a diffraction pattern or diffractive optical element (DOE) is incorporated within or imprinted / embossed on a light guiding optical element (LOE, e.g., a planar waveguide) such that as collimated light (a light beam with a substantially planar wavefront) is substantially totally internally reflected along the LOE, it intersects the diffraction pattern at multiple locations and exits towards the user's eye. The DOE is configured such that the light exiting through it from the LOE is converged and appears to arise from a particular depth plane. The collimated light may be generated using an optical condenser lens ("condenser").
[0059] For example, the first LOE may be configured to deliver collimated light to the eye such that it appears to originate from an optically infinite depth plane (0 diopters). Another LOE may be configured to deliver collimated light such that it appears to originate from a distance of 2 meters (1 / 2 diopter). Yet another LOE may be configured to deliver collimated light such that it appears to originate from a distance of 1 meter (1 diopter). By using a stacked LOE assembly, it can be understood that multiple depth planes can be created and each LOE is configured to display an image that appears to originate from a particular depth plane. It should be understood that the stack may include any number of LOEs. However, at least N stacked LOEs are required to generate N depth planes. Further, N, 2N, or 3N stacked LOEs may be used to generate an RGB color image on N depth planes.
[0060] To present 3D virtual content to the user, VR, AR, and / or MR systems project an image of the virtual content into the user's eye such that it appears to originate from various depth planes in either blend or discrete mode in the Z direction (i.e., orthogonal to recede from the user's eye). In other words, the virtual content can vary not only in the X and Y directions (i.e., in a 2D plane orthogonal to the user's eye's central visual axis), but also in the Z direction such that the user can perceive the object as being at a very close distance or at an infinite distance or at any distance in between. In other embodiments, the user can simultaneously perceive multiple objects at different depth planes. For example, the user may be able to simultaneously see a virtual bird at a distance of 3 meters from the user and a virtual coffee cup at arm's length (about 1 meter) from the user. Alternatively, the user may be able to see a virtual dragon appear from infinity and run towards the user.
[0061] The multi-plane focus system creates a perception of variable depth by projecting an image onto some or all of a plurality of depth planes located at an individual fixed distance in the Z direction from the user's eyes, either in blend or discrete mode. Referring now to FIG. 6, it should be understood that the multi-plane focus system may display the frame on a fixed depth plane 150 (e.g., the six depth planes 150 shown in FIG. 6). The MR system can include any number of depth planes 150, but one exemplary multi-plane focus system has six fixed depth planes 150 in the Z direction. When generating virtual content on one or more of the six depth planes 150, a 3D perception is created such that the user perceives one or more virtual objects at a variable distance from the user's eyes. Assuming that the human eye is more sensitive to objects that are closer in distance than objects that appear to be farther away, more depth planes 150 are generated closer to the eye, as shown in FIG. 6. In other embodiments, the depth planes 150 may be equidistantly spaced from each other.
[0062] The depth plane position 150 may be measured in diopters, which is a unit of refractive power equal to the reciprocal of the focal length measured in meters. For example, in some embodiments, the depth plane DP1 may be 1 / 3 diopter apart, the depth plane DP2 may be 0.3 diopter apart, the depth plane DP3 may be 0.2 diopter apart, the depth plane DP4 may be 0.15 diopter apart, the depth plane DP5 may be 0.1 diopter apart, and the depth plane DP6 may represent infinity (i.e., 0 diopters apart). It should be understood that other embodiments may generate the depth plane 150 at other distances / diopters. Thus, when generating virtual content at the strategically placed depth plane 150, the user is able to perceive virtual objects in three dimensions. For example, the user may perceive that a first virtual object is close when it is displayed on the depth plane DP1 while another virtual object appears at infinity on the depth plane DP6. Alternatively, the virtual object may first be displayed on the depth plane DP6 and then on the depth plane DP5 etc. until the virtual object appears very close to the user. It should be understood that the above examples are significantly simplified for illustrative purposes. In another embodiment, all six depth planes may be concentrated on a particular focal length away from the user. For example, if the virtual content to be displayed is a coffee cup 0.5 meters away from the user, all six depth planes may be generated at various cross-sections of the coffee cup, giving the user a very granular 3D view of the coffee cup.
[0063] In some embodiments, the VR, AR, and / or MR system may act as a multifocal system. In other words, images may be generated rapidly and continuously such that all six LOEs appear to result from six fixed depth planes, and the light sources may be illuminated almost simultaneously such that the image information is rapidly transmitted to LOE1, then LOE2, then LOE3, etc. For example, a portion of a desired image, including an empty image at optical infinity, may be input at time 1, and an LOE (e.g., depth plane DP6 from FIG. 6) that retains the collimation of light may be utilized. Then, an image of a closer tree branch may be input at time 2, and an LOE configured to create an image that appears to result from a depth plane 10 meters away (e.g., depth plane DP5 from FIG. 6) may be utilized. Then, an image of a pen may be input at time 3, and an LOE configured to create an image that appears to result from a depth plane 1 meter away may be utilized. This type of paradigm can be repeated in a high-speed time-series manner such that the user's eyes and brain (e.g., visual cortex) perceive the inputs as all being part of the same image.
[0064] In blend mode, such a multifocal system has system requirements from the perspective of display and processor speed that may not be achievable in a portable VR, AR, and / or MR system. To overcome these system limitations, the blend mode system may display the images adjacent to the user / viewer's depth of focus more clearly while displaying other images less clearly at a lower sacrifice from the perspective of display and processor speed. In discrete mode, such a multifocal system may project all 3D image content regardless of the optical depth to a single depth plane corresponding to the user / viewer's focus position. However, all discrete mode display systems require identification of the user / viewer's depth of focus.
[0065] A VR, AR, and / or MR system may project an image that appears to originate from various locations along the Z-axis (i.e., the depth plane) to generate an image for a 3D experience / scenario (i.e., by diverging or converging a light beam). As used herein, a light beam includes a directional projection of light energy (including visible and invisible light energy) emitted from a light source. Generating an image that appears to originate from various depth planes conforms to the user's eye's convergence / divergence movement and accommodation for that image, minimizing or eliminating convergence / divergence movement-accommodation conflict.
[0066] Referring now to FIG. 7, an exemplary embodiment of an AR or MR system 700 (hereinafter referred to as "system 700") is illustrated. System 700 uses stacked light guiding optical elements (hereinafter referred to as "LOE 790"). System 700 generally includes one or more image generation processors 710, one or more light sources 720, one or more controller / display bridges (DB) 730, one or more spatial light modulators (SLM) 740, and one or more sets of stacked LOE 790 that function as a multi-focal system. System 700 may also include an eye tracking subsystem 750.
[0067] The image generation processor 710 is configured to generate virtual content to be displayed to the user. The image generation processor 710 may convert an image or video associated with the virtual content into a format that can be projected to the user in 3D. For example, when generating 3D content, the virtual content may need to be formatted such that certain portions of the image are displayed on a particular depth plane while others are displayed on other depth planes. In one embodiment, all of the images may be generated on a particular depth plane. In another embodiment, the image generation processor 710 may be programmed to provide slightly different images to the right and left eyes such that the virtual content appears coherent and comfortable to the user's eyes when viewed together.
[0068] The image generation processor 710 may further include a memory 712, a GPU 714, a CPU 716, and other circuitry for image generation and processing. The image generation processor 710 may be programmed with the desired virtual content to be presented to the user of the system 700. It should be understood that in some embodiments, the image generation processor 710 may be stored within the system 700. In other embodiments, the image generation processor 710 and other circuitry may be stored within a belt pack coupled to the system 700. In some embodiments, the image generation processor 710 or one or more components thereof may be part of a local processing and data module (e.g., local processing and data module 130). As described above, the local processing and data module 130 may be mounted in various configurations, such as fixedly attached to the frame structure 102 (FIG. 2), fixedly attached to a helmet or cap 56 (FIG. 3), removably attached to the torso 58 of the end user 50 (FIG. 4), or removably attached to the waist 60 of the end user 50 in a belt-coupled configuration (FIG. 5).
[0069] The image generation processor 710 is operably coupled to a light source 720 that projects light associated with the desired virtual content and one or more spatial light modulators 740. The light source 720 is compact and has high resolution. The light source 720 is operably coupled to the controller / DB 730. The light source 720 may include color-specific LEDs and lasers arranged in various geometric configurations. Alternatively, the light source 720 may include LEDs or lasers of the same color, each linked to a specific region of the field of view of the display. In another embodiment, the light source 720 may include a broad area emitter such as an incandescent or fluorescent lamp with a mask overlay for segmentation of the emission area and position. The light source 720 is directly connected to the system 700 in FIG. 2B, but the light source 720 may be connected to this system 700 via an optical fiber (not shown). The system 700 may also include a condenser (not shown) configured to collimate the light from the light source 720.
[0070] In various exemplary embodiments, the SLM 740 may be reflective (e.g., LCOS, FLCOS, DLP DMD, or MEMS mirror system), transmissive (e.g., LCD), or emissive (e.g., FSD or OLED). The type of SLM 740 (e.g., speed, size, etc.) can be selected to improve the creation of 3D perception. A DLP DMD operating at a higher refresh rate can be easily incorporated into the stationary system 700, but the wearable system 700 may use a DLP of smaller size and power. The power of the DLP changes the way the 3D depth plane / focus plane is created. The image generation processor 710 is operably coupled to the SLM 740, which encodes the light from the light source 720 with the desired virtual content. The light from the light source 720 may be encoded with image information when it reflects from, emits from, or passes through the SLM 740.
[0071] Light from the SLM740 is directed to the LOE790 such that a light beam encoded with image data by the SLM740 for one depth plane and / or color is substantially propagated along a single LOE790 for delivery to the user's eye. Each LOE790 is configured to project onto the user's retina an image or sub-image that appears to originate from a desired depth plane or FOV angular position. The light source 720 and the LOE790 can thus selectively project images (encoded by the SLM740 under the control of the controller / DB730 in synchronization) that appear to originate from various depth planes or positions within the space. Using each of the light source 720 and the LOE790 to sequentially project images at a sufficiently high frame rate (e.g., 360 Hz for six depth planes at a virtually complete volume frame rate of 60 Hz), the system 700 can generate 3D images of virtual objects that appear to co-exist within the 3D image at various depth planes.
[0072] The controller / DB730 communicates with and is operatively coupled to the image generation processor 710, the light source 720, and the SLM740, and coordinates the synchronous display of images by instructing the SLM740 to encode the light beam from the light source 720 with appropriate image information from the image generation processor 710. The system includes the image generation processor 710, but in some embodiments, the controller / DB730 may also perform at least some of the image generation processes, including, for example, the processes of the memory 712, the GPU 714, and / or the CPU 716. In some embodiments, the controller / DB730 may include one or more components shown within the image generation processor 710, such as, for example, the memory 712, the GPU 714, and / or the CPU 716.
[0073] System 700 may also include an eye-tracking subsystem 750 configured to track the user's eyes and determine the user's convergence / divergence motion points and the direction for each of the user's eyes. The eye-tracking subsystem 750 may also be configured to identify a depth plane corresponding to the focus of the user / viewer in order to more efficiently generate an image and / or display the image in a discrete mode (e.g., using the user's convergence / divergence motion points and / or the direction for each of the user's eyes). In one embodiment, the system 700 is configured to illuminate a subset of the LOEs 790 in discrete mode based on an input from the eye-tracking subsystem 750 such that the image is generated on a desired depth plane that coincides with the focus of the user, as shown in FIGS. 9A and 9B. For example, when the focus of the user is at infinity (e.g., the user's eyes are parallel to each other and other factors are as described below), the system 700 may illuminate the LOEs 790 configured to deliver collimated light to the user's eyes such that the image appears to originate from optical infinity. In another example, if the eye-tracking subsystem 750 determines that the focus of the user is 1 meter away, the LOEs 790, configured to be approximately in focus within that range, may instead be illuminated. In an embodiment where depth planes approximating the entire depth range are illuminated simultaneously (or nearly simultaneously) in blend mode, the eye-tracking subsystem 750 may still enable the rendering of a more efficient image by identifying the focus of the user. Exemplary User / Viewer Focus Depth Determination System and Method
[0074] FIG. 12A depicts a discrete mode display in which the near and far cubes 810, 812, and cylinder 814 are projected onto the near depth plane 0 in response to tracking the viewer 816's focus on the near cube 810 adjacent to the near depth plane 0. The projection onto the near depth plane 0 provides the viewer with depth accommodation to the near depth plane 0 even with respect to the far cube 812 and cylinder 814 that are closer to the far depth plane 1 than the near depth plane 0. To operate in discrete mode, the system must identify or obtain the depth plane corresponding to the focus of the viewer 816 (the "focus plane").
[0075] Figures 13 and 14 depict the system, and Figure 16 depicts a method for determining the focal plane of viewer 816. At step 1612 (see Figure 16), the display system (e.g., its eye-tracking subsystem) tracks the line-of-sight path 1350 of viewer 816 (see Figures 13 and 14). The display system may use an eye-tracking subsystem (described above) to track the line-of-sight path 1350. In some embodiments, the display system tracks one or more line-of-sight paths. For example, the display system may track the line-of-sight path for only the left eye, only the right eye, or both eyes of viewer 816. At step 1614, the display system analyzes 3D image data to identify one or more virtual objects along the line-of-sight path of viewer 816. The step of analyzing 3D image data to identify one or more virtual objects along the line-of-sight path of viewer 816 may include the step of analyzing depth segmentation data corresponding to the 3D data. If only one virtual object intersects the line-of-sight path of viewer 816, as determined at step 1616, the display system identifies the depth of the intersecting virtual object at the intersection as the depth of the focal plane of viewer 816 at step 1618.
[0076] In particular, FIGS. 13 and 14 show that the line of sight path 1350 of viewer 816 intersects only the near cube 810. FIG. 14 shows that the near cube 810 has a depth of 1410, the far cube 812 has a depth of 1412, and the cylinder 814 has a depth of 1414. FIG. 14 also shows that the system is capable of projecting an image at one or more depths 1408. Since the system has determined that the near cube 810 is the only virtual object that intersects the line of sight path 1305 of viewer 816, the system identifies the depth 1410 of the near cube 810 as the depth of the focal plane of viewer 816. Although not necessary, the system and method for determining the focal plane of viewer 816 is more accurate if the virtual objects 810, 812, 814 are already displayed to viewer 816. In such an embodiment, the line of sight path 1350 of viewer 816 indicates that viewer 816 is likely to focus on a particular virtual object (e.g., the near cube 810).
[0077] In embodiments where the viewer's line of sight path intersects more than one virtual object (i.e., "no" in step 1616), the display system may proceed along an optional method in steps 1620-1626. In step 1620, the display system tracks the viewer's second line of sight path. In some embodiments, the first line of sight path is associated with the user's first eye and the second line of sight path is associated with the user's second eye. In step 1622, the display system analyzes the 3D image data to identify one or more virtual objects along the viewer's second line of sight path. If there is only one virtual object that intersects both the viewer's line of sight path and the second line of sight path, as determined in step 1624, the display system, in step 1626, identifies the depth of the intersecting virtual object as the depth of the viewer's focal plane.
[0078] In some embodiments, the display system may use an eye-tracking subsystem to determine the vergence / accommodation motion point of the viewer's eyes. The display system may then use the determined vergence / accommodation motion point to confirm that identification of the depth of the viewer's focal plane.
[0079] In discrete display mode embodiments, the display system may use the viewer's identified focal plane (e.g., its depth) to display virtual objects corresponding to 3D image data.
[0080] FIG. 12B depicts a blended-mode composite light field according to some embodiments. In the blend display mode, the near and far cubes 810, 812, and the cylinder 814 are "virtually projected" onto individual virtual depth planes 1210, 1212, 1214 adjacent to the individual locations of the near and far cubes 810, 812 and the cylinder 814 along the optical axis ("Z-axis") of the system. Only the near and far depth planes 0, 1 are actually illuminated, so only the near and far cubes 810, 812, and the cylinder 814 are virtually projected. As described above, the virtual projection onto the virtual depth planes 1210, 1212, 1214 is simulated through a scalar blend of the sub-images projected onto the near and far depth planes 0, 1. The rendering of the 2D images projected onto each of the virtual depth planes 1210, 1212, 1214 may include blending of the objects therein.
[0081] Identification of the viewer's focal plane (e.g., its depth) is not necessary to display an object in the blend display mode, but identifying the viewer's focal plane enables the blend-mode display system to operate within its hardware limitations (processing, display, memory, communication channels, etc.). In particular, when the viewer's focal plane enables the blend-mode display system to efficiently allocate system resources for the display of virtual objects at the viewer's focal plane.
[0082] Regarding a blended display mode system, the present method can be simplified in that, instead of analyzing 3D data and identifying one or more virtual objects along the line of sight of viewer 816, the display system can analyze only the 2D image frames of the blended mode display to identify one or more virtual objects along the line of sight of viewer 816. This allows the display system to analyze less data, thereby increasing system efficiency and reducing system resource requirements.
[0083] The above embodiments include various different methods for determining the user's depth / viewer's focus. The step of determining the user / viewer's depth of focus enables the display system to project objects in a discrete display mode at the depth of focus, thereby reducing the requirements on the display system. Such reduced requirements include display and processor speed, power, and heat generation. The step of determining the user / viewer's depth of focus also enables the display system to project objects in a blended display mode while allocating system resources to the image on the user / viewer's focal plane. Further, the user / viewer focal plane determination step described above is more efficient than other focal plane determination techniques and improves processor performance in terms of speed, power, and heat generation. System Architecture Overview
[0084] FIG. 17 is a block diagram of an exemplary computing system 1700 suitable for implementing embodiments of the present disclosure. Computer system 1700 includes a bus 1706 or other communication mechanism for communicating information, which interconnects a processor 1707, a system memory 1708 (e.g., RAM), a static storage device 1709 (e.g., ROM), a disk drive 1710 (e.g., magnetic or optical), a communication interface 1714 (e.g., a modem or Ethernet card), a display 1711 (e.g., a CRT or LCD), an input device 1712 (e.g., a keyboard), and subsystems and devices such as cursor control.
[0085] According to one embodiment of the present disclosure, computer system 1700 performs specific operations by the processor 1707 executing one or more sequences of one or more instructions contained within the system memory 1708. Such instructions may be read into the system memory 1708 from another computer-readable / usable medium, such as the static storage device 1709 or the disk drive 1710. In alternative embodiments, wire circuitry may be used in place of or in combination with software instructions to implement the present disclosure. Accordingly, embodiments of the present disclosure are not limited to any specific combination of hardware circuitry and / or software. In one embodiment, the term "logic" shall mean any combination of software or hardware used to implement all or part of the present disclosure.
[0086] As used herein, the term "computer-readable medium" or "computer-usable medium" refers to any medium involved in providing instructions to the processor 1707 for execution. Such a medium may take many forms, including but not limited to non-volatile media and volatile media. Non-volatile media includes, for example, optical or magnetic disks such as the disk drive 1710. Volatile media includes dynamic memory such as the system memory 1708.
[0087] A computer-readable medium in general form includes, for example, a floppy (registered trademark) disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, any other optical medium, a punch card, a paper tape, any other physical medium with a pattern of holes, a RAM, a PROM, an EPROM, a FLASH-EPROM (e.g., NAND flash, NOR flash), any other memory chip or cartridge, or any other medium from which a computer can read.
[0088] In certain embodiments of the present disclosure, the execution of a sequence of instructions for practicing the present disclosure is performed by a single computer system 1700. According to other embodiments of the present disclosure, two or more computer systems 1700 coupled by a communication link 1715 (e.g., a LAN, a PTSN, or a wireless network) may cooperate with each other to perform a sequence of instructions required to practice the present disclosure.
[0089] The computer system 1700 may transmit and receive programs, i.e., application code, messages, data, and instructions, through the communication link 1715 and the communication interface 1714. The received program code may be executed by the processor 1707 as it is received and / or stored in the disk drive 1710 or other non-volatile storage device for later execution. The database 1732 in the storage medium 1731 may be used to store data accessible by the system 1700 via the data interface 1733.
[0090] The blend mode embodiments described above include two depth planes (e.g., near and far depth planes), although other embodiments may include more than two depth planes. Increasing the number of depth planes increases the fidelity with which virtual depth planes are simulated using them. However, this increase in fidelity is offset by an increase in hardware and processing requirements that may exceed the resources available in current portable 3D rendering and display systems.
[0091] Certain aspects, advantages, and features of the present disclosure are described herein. It should be understood that not necessarily all such advantages can be achieved in accordance with any particular embodiment of the present disclosure. Accordingly, the present disclosure may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.
[0092] Embodiments are described in relation to the accompanying drawings. However, it should be understood that the figures are not drawn to scale. Distances, angles, etc. are merely illustrative and do not necessarily maintain an exact relationship to the actual dimensions and layout of the devices shown. Additionally, the foregoing embodiments are described at a level of detail that enables one of ordinary skill in the art to make and use the devices, systems, methods, and equivalents described herein. Various modifications are also contemplated. Components, elements, and / or steps may be modified, added, removed, or rearranged.
[0093] The devices and methods described herein can advantageously be implemented, at least in part, using, for example, computer software, hardware, firmware, or any combination of software, hardware, and firmware. A software module can include computer-executable code stored within a computer's memory for performing the functions described herein. In some embodiments, the computer-executable code is executed by one or more general-purpose computers. However, one of ordinary skill in the art will understand that, in light of this disclosure, any module that is implemented using software that will be executed on a general-purpose computer can also be implemented using different combinations of hardware, software, or firmware. For example, such a module can be implemented entirely in hardware using combinations of integrated circuits. Alternatively, or in addition, such a module can be implemented using a special-purpose computer designed to perform the specific functions described herein rather than by a general-purpose computer. Additionally, when a method that is performed, or can be performed, at least in part by computer software is described, it should be understood that such a method can be provided on a non-transitory computer-readable medium that, when read by a computer or other processing device, causes the method to be performed.
[0094] While certain embodiments are explicitly described, other embodiments will be apparent to one of ordinary skill in the art based on this disclosure.
[0095] The various processors and other electronic components described herein are suitable for use in combination with any optical system for projecting light. The various processors and other electronic components described herein are also suitable for use in combination with any audio system for receiving voice commands.
[0096] Various exemplary embodiments of the present disclosure are described herein. These examples are referred to in a non-limiting sense. They are provided to illustrate more broadly applicable aspects of the present disclosure. Various changes may be made to the present disclosure described, and equivalents may be substituted without departing from the true spirit and scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation, material, composition, process, process act, or step to the purpose, spirit, or scope of the present disclosure. Further, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein can be readily separated from or combined with features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. All such modifications are intended to be within the scope of the claims associated with the present disclosure.
[0097] The present disclosure includes methods that can be implemented using the devices of the subject matter. The method may include the act of providing such a suitable device. Such providing may be performed by an end user. In other words, the act of "providing" simply requires that the end user act to obtain, access, approach, locate, configure, activate, power on, or otherwise provide the device required in the method of the subject matter. The methods recited herein may be performed in any order of the recited logical possible events and in the recited order of events.
[0098] Exemplary aspects of the present disclosure are described above, along with details regarding material selection and manufacturing. Regarding other details of the present disclosure, these are understood in relation to the patents and publications referenced above and are generally known or understandable by those skilled in the art. The same may apply to the method-based aspects of the present disclosure from the perspective of additional acts generally or logically employed.
[0099] In addition, although the present disclosure has been described with reference to several embodiments that optionally incorporate various features, the present disclosure is not limited to what is described or illustrated as being considered with respect to each variation of the disclosure. Various modifications may be made to the present disclosure as described, and equivalents (whether listed herein or not for purposes of some brevity) may be substituted without departing from the true spirit and scope of the present disclosure. In addition, when a range of values is provided, it is to be understood that all intervening values, as well as any other stated value or intervening values within the stated range between the upper and lower limits of that range, are included within the present disclosure.
[0100] Also contemplated is that any optional feature of the variations described may be claimed and described independently or in combination with any one or more of the features described herein. References to singular items include the possibility that multiple identical items exist. More specifically, as used in this specification and the claims associated herewith, the singular forms “a,” “an,” “said,” and “the” include plural references unless specifically stated otherwise. In other words, the use of the article enables “at least one” of the items of the subject matter in the above description and the claims associated with the present disclosure. Further, it should be noted that such claims may be drafted to exclude any optional element. Accordingly, the recitation herein is intended to serve as a antecedent for the use of exclusive terminology such as “solely,” “only,” and equivalents, and the use of “negative” limitations in connection with the recitation of elements of a claim.
[0101] Without using such exclusive terminology, the term "comprising" in the claims associated with the present disclosure shall be taken to allow the inclusion of any additional elements, whether or not a given number of elements are recited in such claims, or the addition of features may be regarded as transforming the nature of the elements recited in such claims. Unless specifically defined herein, all technical and scientific terms used herein should be given the broadest generally understood meaning possible while maintaining the validity of the claims.
[0102] The scope of the present disclosure should not be limited to the provided examples and / or this specification, but rather should be limited only by the scope of the language of the claims associated with the present disclosure.
[0103] In the foregoing specification, the present disclosure has been described with reference to its specific embodiments. However, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the present disclosure. For example, the foregoing process flow is described with reference to a particular order of process actions. However, many of the orders of the described process actions may be changed without affecting the scope or operation of the present disclosure. The specification and drawings are, therefore, to be regarded in an illustrative rather than a limiting sense.
Claims
【Claim 1】 A method for determining the depth of focus of a user of a three-dimensional (“3D”) display device, comprising: tracking a first line of sight of the user, the first line of sight corresponding to the first eye of the user and not corresponding to the second eye of the user; analyzing 3D data based on the first line of sight and identifying one or more virtual objects along the first line of sight of the user; based on the identified one or more virtual objects and the first line of sight, when it is determined that only one virtual object intersects the first line of sight of the user, analyzing the 3D data to determine the depth of the only one virtual object; determining that the depth of focus of the user is the determined depth of the only one virtual object using only the 3D data for the first line of sight of the user and the only one virtual object; when more than one virtual object intersects the first line of sight of the user, tracking a second line of sight of the user, the second line of sight corresponding to the second eye of the user and not corresponding to the first eye of the user; analyzing the 3D data based on the second line of sight and identifying one or more virtual objects along the second line of sight of the user; based on the identified one or more virtual objects and the first and second lines of sight, when it is determined that only one virtual object intersects both the first line of sight and the second line of sight of the user, analyzing the 3D data to determine the depth of the only one virtual object; determining that the depth of focus of the user is the determined depth of the only one virtual object; and A method comprising **Claim 2** The method according to claim 1, wherein determining the depth of the only one virtual object comprises analyzing the 3D data. **Claim 3** A method for determining the depth of focus of a user of a three-dimensional ( "3D") display device, tracking the first line of sight of the user, the first line of sight corresponding to the first eye of the user and not corresponding to the second eye of the user, analyzing a plurality of two-dimensional ( "2D") image frames displayed by the display device based on the first line of sight of the user, and identifying one or more virtual objects along the first line of sight of the user, when, based on the identified one or more virtual objects and the first line of sight, it is determined that only one of the plurality of 2D image frames contains a virtual object along the first line of sight of the user, analyzing the plurality of 2D image frames and determining the determination of the depth of the only one 2D image frame, determining that the depth of focus of the user is the determined depth of the only one 2D image frame using only the 3D data for the first line of sight of the user and the virtual object comprising **Claim 4** The method according to claim 3, wherein each of the plurality of 2D image frames is displayed at an individual depth from the user. **Claim 5** The method according to claim 3, wherein analyzing the plurality of 2D image frames comprises analyzing depth segmentation data corresponding to the plurality of 2D image frames. **Claim 6** A method for determining the depth of focus of a user of a three-dimensional ( "3D") display device, Tracking the first line-of-sight path of the user, where the first line-of-sight path corresponds to the first eye of the user and does not correspond to the second eye of the user, Analyzing a plurality of two-dimensional ("2D") image frames displayed by the display device based on the first line-of-sight path, and identifying one or more virtual objects along the first line-of-sight path of the user, Based on the identified one or more virtual objects and the first line-of-sight path, when it is determined that only one of the plurality of 2D image frames contains a virtual object along the first line-of-sight path of the user, Analyzing the plurality of 2D image frames and determining the depth determination of the only one 2D image frame, Using only the 3D data about the first line-of-sight path of the user and the virtual object to determine that the focus depth of the user is the determined depth of the only one 2D image frame, When more than one of the plurality of 2D image frames contains a virtual object along the first line-of-sight path of the user, Tracking the second line-of-sight path of the user, where the second line-of-sight path corresponds to the second eye of the user and does not correspond to the first eye of the user, Analyzing the plurality of 2D image frames displayed by the display device based on the second line-of-sight path, and identifying one or more virtual objects along the second line-of-sight path of the user, Based on the identified one or more virtual objects and the first and second line-of-sight paths, when it is determined that only one of the plurality of 2D image frames contains a virtual object along both the first line-of-sight path and the second line-of-sight path of the user, Analyzing the 3D data and determining the depth of the only one 2D image frame, Determining that the depth of focus of the user is the determined depth of the only one 2D image frame A method comprising: **Claim 7** The method according to claim 3, wherein determining the depth of the only one 2D image frame includes analyzing the 2D image frame. **Claim 8** The method according to claim 6, wherein determining the depth of the only one 2D image frame includes analyzing the 2D image frame.
Citation Information
Patent Citations
Display device and display control program
JP2014219621A
Selection of virtual objects in 3D space
JP2018534687A