Continuous Time Warping and Binocular Time Warping and Method for Virtual Reality and Augmented Reality Display Systems
By employing continuous and binocular time warping techniques to transform image frames based on user head movement, the VR and AR systems address latency issues, enhancing the realism and responsiveness of the experiences.
Patent Information
- Application Number
- JP2024023581
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-08-26
- Filing Date
- 2024-02-20
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2037-08-25
AI Technical Summary
Existing VR and AR systems face challenges in providing a seamless and realistic experience due to latency issues caused by user head movement, particularly in scanning displays where the time between the first and last pixel can extend up to the full frame duration, leading to pose inconsistency and undesirable user experiences.
The implementation of continuous and binocular time warping methods that account for user head movement, minimizing latency by transforming image frames based on updated viewer positions, and using external or internal hardware to perform these transformations before final image display.
This approach significantly reduces the latency from motion to photon, enhancing the responsiveness and immersion of VR and AR experiences, and prevents issues like image cracking and pose inconsistency, thereby improving user satisfaction.
Smart Images

Figure 0007683065000001 
Figure 0007683065000002 
Figure 0007683065000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Patent Application No. 62 / 380,302, filed Aug. 26, 2016, entitled “Time Warp for Virtual and Augmented Reality Display Systems and Methods,” which is hereby incorporated by reference in its entirety for all purposes. This application incorporates by reference in its entirety each of the following U.S. patent applications: U.S. Provisional Application No. 62 / 313,698, filed Mar. 25, 2016 (Attorney Docket No. MLEAP.058PR1); U.S. Patent Application No. 14 / 331,218, filed Jul. 14, 2014; U.S. Patent Application No. 14 / 555,585, filed Nov. 27, 2014; U.S. Patent Application No. 14 / 690,401, filed Apr. 18, 2015; U.S. Patent Application No. 14 / 726,424, filed May 29, 2015; U.S. Patent Application No. 14 / 726,429, filed May 29, 2015; U.S. Patent Application No. 15 / 146,296, filed May 4, 2016; U.S. Patent Application No. 15 / 182,511, filed Jun. 14, 2016; U.S. Patent Application No. 15 / 182,528, filed Jun. 14, 2016; U.S. Provisional Application No. 62 / 206,765, filed Aug. 18, 2015 (Attorney Docket No. MLEAP.002PR); U.S. Patent Application No. 15 / 239,710, filed Aug. 18, 2016 (Attorney Docket No. MLEAP.002A); U.S. Provisional Application No. 62 / 377,831, filed Aug. 22, 2016 (Attorney Docket No. 101782 - 1021207(000100US)), and U.S. Provisional Application No. 62 / 380,302, filed Aug. 26, 2016 (Attorney Docket No. 101782 - 1022084(000200US)).
[0002] This disclosure relates to virtual reality and augmented reality visualization systems. More specifically, this disclosure relates to methods of continuous time warping and binocular time warping for virtual reality and augmented reality visualization systems.
Background Art
[0003] Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality" (VR) or "augmented reality" (AR) experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or can be perceived as being real. VR scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. AR scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the actual world surrounding the user.
[0004] For example, referring to FIG. 1, an AR scene 4 is depicted, and to a user of AR technology, a real-world park setting 6 characterized by people, trees, buildings in the background, and a concrete platform 8 is visible. In addition to these items, a user of AR technology also "sees" and perceives as being present a robotic figure 10 standing on the real-world concrete platform 8 and an avatar character 2 in the form of a flying cartoon that appears anthropomorphic like a bumblebee, although these elements (e.g., avatar character 2 and robotic figure 10) do not exist in the real world. Due to the extreme complexity of human visual perception and the nervous system, the production of VR or AR technology that facilitates the comfortable and natural-seeming presentation of virtual image elements among other virtual or real-world image elements is difficult.
[0005] One major problem is directed to modifying virtual images presented to a user based on user movement. For example, when a user moves their head, the visual area (e.g., the field of view) and the viewpoints of objects within the visual area can change. The overlay of content that would be presented to the user needs to be modified in real-time or near real-time to account for user movement and provide a more realistic VR or AR experience.
[0006] The refresh rate of a system determines the rate at which the system generates content and displays (or transmits for display) the generated content to the user. For example, if the refresh rate of the system is 60 Hertz, the system generates content (e.g., renders, modifies, and the like) every 16 milliseconds and displays the generated content to the user. VR and AR systems may generate content based on the user's pose. For example, the system may determine the user's pose within the entire 16-millisecond time window, generate content based on the determined pose, and display the generated content to the user. The time between when the system determines the user's pose and when the system displays the generated content to the user is known as the "motion-to-photon latency". The user may change their pose during the time between when the system determines the user's pose and when the system displays the generated content. If this change is not taken into account, it may result in an undesirable user experience. For example, the system may determine the user's first pose and start generating content based on the first pose. The user may then change their pose to a second pose during the time between when the system determines the first pose and subsequently generates content based on the first pose and when the system displays the generated content to the user. Since the content is generated based on the first pose and the user currently has the second pose, the generated content displayed to the user will appear misaligned to the user due to the pose inconsistency. Pose inconsistency can lead to an undesirable user experience.
[0007] As a post - processing step that acts on the buffered image, for example, corrections may be applied across the entire rendered image frame, taking into account that the user has changed their user pose. This technique can work for panel displays that display an image frame by flushing / lighting all pixels (e.g., within 2 ms) when all pixels are rendered, but this technique may not function well in a scanning display that displays the image frame pixel - by - pixel in a sequential manner (e.g., within 16 ms). In a scanning display that displays the image frame pixel - by - pixel in a sequential manner, the time between the first pixel and the last pixel can extend up to the full frame duration (e.g., 16 ms for a 60 Hz display), during which the user pose can be significantly changed.
[0008] Embodiments address these and other problems associated with VR or AR systems that implement conventional time warping.
Summary of the Invention
Means for Solving the Problems
[0009] The present disclosure relates to techniques that enable three - dimensional (3D) visualization systems. More specifically, the present disclosure addresses components, sub - components, architectures, and systems for producing augmented reality (「AR」) content for a user through a display system that enables the perception of virtual reality (「VR」) or AR content as if it were occurring within the observed real world. Such immersive sensory input may also be referred to as mixed reality (「MR」).
[0010] In some embodiments, the light pattern is input into a waveguide of a display system configured to present the content to a user wearing the display system. The light pattern may be input by an optical projector, and the waveguide may be configured to propagate light of a particular wavelength through total internal reflection within the waveguide. The optical projector may include a light emitting diode (LED) and a liquid crystal on silicon (LCOS) system. In some embodiments, the optical projector may include a scanning fiber. The light pattern may include image data in a time-sequence format.
[0011] Various embodiments provide continuous and / or binocular time warping methods that account for a user's head movement and minimize the latency from the motion resulting from the user's head movement to the photons. Continuous time warping enables the conversion of an image from a first viewpoint (e.g., based on a first position of the user's head) to a second viewpoint (e.g., based on a second position of the user's head) without the need to re-render the image from the second viewpoint. In some embodiments, continuous time warping is implemented on external hardware (e.g., a controller external to the display), and in other embodiments, continuous time warping is implemented on internal hardware (e.g., a controller internal to the display). Continuous time warping is performed before the final image is displayed on a display device (e.g., a sequential display device).
[0012] Some embodiments provide a method for transforming an image frame based on an updated position of a viewer. The method may include obtaining, by a computing device, a first image frame from a graphics processing unit. The first image frame corresponds to a first viewpoint associated with a first position of the viewer. The method may also include receiving data associated with a second position of the viewer. The computing device may continuously transform at least a portion of the first image frame pixel-by-pixel to generate a second image frame. The second image frame corresponds to a second viewpoint associated with the second position of the viewer. The computing device may transmit the second image frame to a display module of the head-mounted display device for display thereon.
[0013] Various embodiments provide a method for transforming an image frame based on an updated position of a viewer. The method may include, by a graphics processing unit, at a first time, rendering a left image frame for a left display of a binocular head-mounted display device. The left image frame corresponds to a first viewpoint associated with a first position of the viewer. The method may also include, by a computing device, rendering, from the graphics processing unit, a right image frame for a right display of the binocular head-mounted display device. The right image frame corresponds to the first viewpoint associated with the first position of the viewer. The graphics processing unit may receive, at a second time after the first time, data associated with a second position of the viewer. The data includes a first pose estimate based on the second position of the viewer. The graphics processing unit may use the first pose estimate based on the second position of the viewer to transform at least a portion of the left image frame and generate an updated left image frame for the left display of the binocular head-mounted display device. The updated left image frame corresponds to a second viewpoint associated with the second position of the viewer. The graphics processing unit may transmit, at a third time after the second time, the updated left image frame to the left display of the binocular head-mounted display device for display on the left display. The graphics processing unit may receive, at a fourth time after the second time, data associated with a third position of the viewer. The data includes a second pose estimate based on the third position of the viewer. The graphics processing unit may use the second pose estimate based on the third position of the viewer to transform at least a portion of the right image frame and generate an updated right image frame for the right display of the binocular head-mounted display device. The updated right image frame corresponds to a third viewpoint associated with the third position of the viewer. The graphics processing unit may transmit, at a fifth time after the fourth time, the updated right image frame to the right display of the binocular head-mounted display device for display on the right display.
[0014] An embodiment may include a computing system including at least a graphics processing unit, a controller, and an eyepiece display device to implement the method steps described above.
[0015] Additional features, advantages, and embodiments are described in the following forms, figures, and claims for carrying out the invention. The present invention provides, for example, the following. (Item 1) A method for converting an image frame based on an updated position of a viewer, the method comprising: acquiring, by a computing device, from a graphics processing unit, a first image frame corresponding to a first viewpoint associated with a first position of the viewer; receiving data associated with a second position of the viewer; continuously converting, by the computing device, at least a part of the first image frame pixel by pixel to generate a second image frame corresponding to a second viewpoint associated with the second position of the viewer; transmitting, by the computing device, the second image frame to a display module of the eyepiece display device for display on the eyepiece display device and including. (Item 2) The method according to item 1, wherein the second position corresponds to a rotation around an optical center of the eyepiece display device from the first position. (Item 3) The method according to item 1, wherein the second position corresponds to a horizontal translation from the first position. (Item 4) The method according to item 1, wherein the first image frame is continuously converted until the converted first image frame is converted into photons. (Item 5) The method according to item 1, wherein the first image frame is a subsection of an image streamed from the graphics processing unit. (Item 6) The computing device includes a frame buffer for receiving the first image frame from the graphics processing unit, and the step of continuously converting at least a part of the first image frame pixel by pixel further includes the step of redirecting, by the computing device, display device pixels of the eyepiece display device from default image pixels in the frame buffer to different image pixels in the frame buffer The method according to item 1, including this. (Item 7) The computing device includes a frame buffer for receiving the first image frame from the graphics processing unit, and the step of continuously converting at least a part of the first image frame pixel by pixel further includes the step of transmitting, by the computing device, image pixels in the frame buffer to display device pixels of the eyepiece display device that are different from default display device pixels assigned to the image pixels The method according to item 1, including this. (Item 8) The computing device includes a frame buffer for receiving the first image frame from the graphics processing unit, and the step of continuously converting further includes the step of receiving the first image frame from the frame buffer of the graphics processing unit while converting at least a part of the first image frame including, and the step of converting is A step of receiving, by the computing device, a first image pixel in a frame buffer of the graphics processing unit as a first image pixel in a frame buffer of the computing device, wherein the first image pixel in the frame buffer of the graphics processing unit is first assigned to a second image pixel in the frame buffer of the computing device The method according to item 1, including (Item 9) The data associated with the second position of the viewer includes optical data and data from an inertial measurement unit, and the computing device includes an attitude estimator module for receiving the optical data and the data from the inertial measurement unit, and a frame buffer for receiving the first image frame from the graphics processing unit. The method according to item 1 (Item 10) The computing device includes a frame buffer for receiving the first image frame from the graphics processing unit, and an external buffer processor for converting at least a part of the first image frame into the second image frame by shifting buffered image pixels stored in the frame buffer prior to transmitting the second image frame to the eyepiece display device. The method according to item 1 (Item 11) The first image pixel in the frame buffer is shifted by a first amount, and the second image pixel in the frame buffer is shifted by a second amount different from the first amount. The method according to item 10 (Item 12) A method for converting an image frame based on an updated position of a viewer, the method comprising A step of rendering, by a graphics processing unit, a left image frame for a left display of a binocular head-mounted display device at a first time, wherein the left image frame corresponds to a first viewpoint associated with a first position of the viewer. A step of rendering, by a computing device, a right image frame for a right display of the binocular head-mounted display device from the graphics processing unit, wherein the right image frame corresponds to a first viewpoint associated with a first position of the viewer. A step of receiving, by the graphics processing unit, data associated with a second position of the viewer at a second time after the first time, wherein the data includes a first pose estimation based on the second position of the viewer. A step of using, by the graphics processing unit, a first pose estimation based on the second position of the viewer to transform at least a part of the left image frame and generate an updated left image frame for the left display of the binocular head-mounted display device, wherein the updated left image frame corresponds to a second viewpoint associated with the second position of the viewer. A step of transmitting, by the graphics processing unit, the updated left image frame to the left display of the binocular head-mounted display device to be displayed on the left display at a third time after the second time. A step of receiving, by the graphics processing unit, data associated with a third position of the viewer at a fourth time after the second time, wherein the data includes a second pose estimation based on the third position of the viewer. A step of using a second pose estimation based on a third position of the viewer by the graphic processing unit to convert at least a part of the right image frame and generate an updated right image frame for a right display of the binocular head-mounted display device, wherein the updated right image frame corresponds to a third viewpoint associated with the third position of the viewer. A step of transmitting the updated right image frame to a right display of the binocular head-mounted display device by the graphic processing unit so as to be displayed on the right display at a fifth time after the fourth time. A method comprising the above. (Item 13) The method according to item 12, wherein the right image frame is rendered by the graphic processing unit at the first time. (Item 14) The method according to item 12, wherein the right image frame is rendered by the graphic processing unit at a sixth time after the first time and before the fourth time. (Item 15) The method according to item 12, wherein the rendered right image frame is the same as the left image frame rendered at the first time. (Item 16) A system comprising: A graphic processing unit configured to generate a first image frame corresponding to a first viewpoint associated with a first position of a viewer. A controller, wherein the controller: Receives data associated with a second position of the viewer. Continuously converts at least a part of the first image frame pixel by pixel and generates a second image frame corresponding to a second viewpoint associated with the second position of the viewer. Streams the second image frame. A controller configured to perform; An eyepiece display device, wherein the eyepiece display device Receives the second image frame by the controller as it is being streamed; Converts the second image frame into photons; Emits the photons towards the viewer An eyepiece display device configured to perform; and A system comprising. (Item 17) The system according to item 16, wherein the eyepiece display device is a sequential scanning display device. (Item 18) The system according to item 16, wherein the controller is incorporated within the eyepiece display device. (Item 19) The system according to item 16, wherein the controller is coupled to and provided between the graphics processing unit and the eyepiece display device. (Item 20) The controller further comprises A frame buffer for receiving the first image frame from the graphics processing unit; and An external buffer processor for converting at least a portion of the first image frame into the second image frame by offsetting buffered image pixels stored in the frame buffer prior to streaming the second image frame to the eyepiece display device The system according to item 16, comprising.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
DETAILED DESCRIPTION OF THE INVENTION
[0017] A virtual reality (“VR”) experience may be provided to a user through a wearable display system. FIG. 2 illustrates an example of a wearable display system 80 (hereinafter referred to as “system 80”). System 80 includes a head-mounted display device 62 (hereinafter referred to as “display device 62”) and various mechanical and electronic modules and systems for supporting the functions of display device 62. Display device 62 may be coupled to a frame 64, which is wearable by a user or viewer 60 (hereinafter referred to as “user 60”) of the display system and is configured to position display device 62 in front of the eyes of user 60. According to various embodiments, display device 62 may be a sequential display. Display device 62 may be for monocular or binocular use. In some embodiments, a speaker 66 is coupled to frame 64 and positioned proximate to the ear canal of user 60. In some embodiments, another speaker (not shown) is positioned adjacent to the other ear canal of user 60 to provide stereo / formable sound control. Display device 62 is operably coupled 68 to a local data processing module 70 by means such as a wired conductor or wireless connectivity, which may be mounted in various configurations, such as fixed to frame 64, fixed to a helmet or hat worn by user 60, built into headphones, or alternatively removably attached to user 60 (e.g., in a backpack configuration, a belt-coupled configuration).
[0018] The local data processing module 70 may include a processor and a digital memory such as a non-volatile memory (e.g., flash memory), both of which can be utilized to assist in data processing, caching, and storage. The data includes a) data captured from sensors such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope (e.g., may be operably coupled to the frame 64 or otherwise attached to the user 60), and / or b) potentially data obtained and / or processed using the remote processing module 72 and / or the remote data repository 74 for passage to the display device 62 after processing or reading. The local data processing module 70 may be operably coupled to the remote processing module 72 and the remote data repository 74, respectively, via communication links 76, 78, such as a wired or wireless communication link, so that these remote modules 72, 74 are operably coupled to each other and available as resources to the local processing and data module 70.
[0019] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the remote data repository 74 may include a digital data storage facility, which may be available through other networking configurations in an Internet or "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local data processing module 70, enabling fully autonomous use from the remote modules.
[0020] In some embodiments, the local data processing module 70 is operably coupled to the battery 82. In some embodiments, the battery 82 is a removable power source such as a commercially available battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 allows the user 60 to operate the system 80 over a longer time period without being connected to a power source and without having to charge a lithium-ion battery or disconnect the system 80 and replace the battery, and includes both an internal lithium-ion battery that can be charged by the user 60 during non-operation of the system 80 and a removable battery.
[0021] FIG. 3A illustrates a user 30 wearing an augmented reality (“AR”) display system that renders AR content as the user 30 moves through a real-world environment 32 (hereinafter referred to as “environment 32”). The user 30 positions the AR display system at location 34, and the AR display system records ambient information of the passable world with respect to location 34, such as mapped features or postures associated with directional audio inputs (e.g., digital representations of objects in the real world that can be memorized and updated as the objects in the real world change). Location 34 is aggregated and processed into data input 36 by at least a passable world module 38 such as the remote processing module 72 of FIG. 2. The passable world module 38 determines where and how the AR content 40 can be placed in the real world, such as on a fixed element 42 (e.g., a table), or in a structure not yet in the field of view 44, or with respect to a mapped mesh model 46 of the real world. As depicted, the fixed element 42 serves as a substitute for any fixed element in the real world, which may be stored in the passable world module 38 so that the user 30 can perceive the content on the fixed element 42 without having to map the fixed element 42 each time it is visible to the user 30. The fixed element 42 may thus be a mapped mesh model from a previous modeling session or determined from a separate user, but nevertheless may be stored on the passable world module 38 for future reference by multiple users. Thus, the passable world module 38 can recognize environment 32 from a previously mapped environment without the user 30's device first mapping environment 32, display the AR content, save the calculation process and cycle, and avoid the latency of any rendered AR content.
[0022] Similarly, a real-world mapped mesh model 46 can be created by the AR display system, interact with the AR content 40, and have appropriate surfaces and metrics for displaying it mapped and stored within the passable world module 38 for future retrieval by the user 30 or other users without the need for remapping or modeling. In some embodiments, data inputs 36 such as geographical location, user identification, and current activity are input to the passable world module 38 to indicate a fixed element 42 of one or more fixed elements available, the AR content 40 last placed on the fixed element 42, and whether to display that same content (such AR content being "persistent" content regardless of whether the user views a particular passable world model).
[0023] FIG. 3B illustrates an overview of the viewing optical assembly 48 and associated components. In some embodiments, two eye tracking cameras 50 oriented towards the user's eyes 49 detect metrics of the user's eyes 49 such as eye shape, eyelid occlusion, pupil direction, and glints on the user's eyes 49. In some embodiments, a depth sensor 51 such as a time-of-flight sensor emits a relay signal into the world and determines the distance to a given object. In some embodiments, a world camera 52 records a wide area peripheral view, maps the environment 32, and detects inputs that can affect the AR content. A camera 53 can further capture a specific timestamp of the real-world image within the user's field of view. The world camera 52, camera 53, and depth sensor 51 each have individual fields of view 54, 55, and 56, respectively, and collect and record data from real-world scenes such as the real-world environment 32 depicted in FIG. 3A.
[0024] The inertial measurement unit 57 may determine the movement and orientation of the visual optical assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 51 is operably coupled to the eye tracking camera 50 as a confirmation of the focusing adjustment measured against the actual distance that the user's eye 49 is looking at.
[0025] In an AR system, when the position of the user 30 changes, the rendered image needs to be adjusted to take into account the new area of the user 30's view. For example, referring to FIG. 2, when the user 60 moves their head, the image displayed on the display device 62 needs to be updated. However, when the user 60's head is in motion and the system 80 needs to determine a new perspective view of the rendered image based on the new head pose, there may be a delay when rendering the image on the display device 62.
[0026] According to various embodiments, the image to be displayed may not need to be re-rendered to save time. Rather, the image may be transformed to match the new perspective of the user 60 (e.g., the new area of the view). This fast image readjustment / view correction may be referred to as time warping. Time warping may enable the system 80 to appear more responsive and immersive even as the head position, and thus the user 60's perspective, changes.
[0027] Time warping may be used to prevent undesirable effects such as cracks on the displayed image. Image cracking is a visual artifact within the display device 62, and the display device 62 shows information from multiple frames in a single screen depiction. Cracking can occur when the frame transmission rate to the display device 62 is not synchronized with the refresh rate of the display device 62.
[0028] Figure 4 illustrates a method by which time warping can be performed once 3D content is rendered. The system 100 illustrated in Figure 4 includes a pose estimator 101 that receives image data 112 and inertial measurement unit (IMU) data 114 from one or more IMUs. The pose estimator 101 may then generate a pose 122 based on the received image data 112 and IMU data 114 and provide the pose 122 to a 3D content generator 102. The 3D content generator 102 may generate 3D content (e.g., 3D image data) and provide the 3D content to a graphics processing unit (GPU) 104 for rendering. The GPU 104 may render the received 3D content at time t1 116 and provide the rendered image 125 to a time warping module 106. The time warping module 106 may receive the rendered image 125 from the GPU 104 and the latest pose 124 from the pose estimator 101 at time t2 117. The time warping module 106 may then perform time warping on the rendered image 125 using the latest pose 124 at time t3 118. The transformed image 126 (i.e., the image on which time warping is performed) is transmitted to a display device 108 (e.g., the display device 62 of Figure 1). Photons are generated at the display device 108 at time t4 120 and emitted towards the user's eye 110, thereby displaying the image on the display device 108. The time warping illustrated in Figure 4 enables the presentation of the latest pose update information (e.g., the latest pose 124) on the image displayed on the display device 108. Older frames (i.e., previously displayed frames or frames received from the GPU) may be used to interpolate the time warping. By using time warping, the latest pose 124 can be incorporated into the displayed image data.
[0029] In some embodiments, the time warping may be parameter warping, non-parameter warping, or asynchronous warping. Parameter warping involves affine operations such as translation, rotation, and scaling of an image. In parameter warping, the pixels of the image are repositioned in a uniform manner. Thus, while parameter warping may be used to correctly update the scene for rotation of the user's head, parameter warping may not account for translation of the user's head, and some regions of the image may be affected differently than others.
[0030] Non-parameter warping involves non-parameter distortion of sections of the image (e.g., stretching of a portion of the image). Non-parameter warping may update the pixels of the image differently in different regions of the image, but non-parameter warping may only partially account for translation of the user's head due to the concept of "disocclusion." Disocclusion may refer to, for example, changes in the user's pose, removal of obstacles in the line of sight, and exposure of objects to the view or reappearance of previously hidden objects from the view as a result of the same.
[0031] Synchronous time warping may refer to warping that separates scene rendering and time warping into two separate asynchronous operations. Asynchronous time warping may be executed on a GPU or external hardware. Asynchronous time warping may increase the frame rate of the displayed image above the rendering rate.
[0032] Time warping according to various embodiments may be performed in response to a new head position of the user (i.e., the attributed head pose). For example, as illustrated in FIG. 4, at time t0 115, the user may move their head (e.g., the user rotates, translates, or both). As a result, the user's viewpoint may change. This will result in a change in what the user sees. Therefore, the rendered image needs to be updated to account for the user's head movement for a realistic VR or AR experience. That is, the rendered image 125 is warped to align (e.g., correspond) with the new head position such that the user perceives the virtual content with the correct spatial positioning and orientation relative to the user's viewpoint within the image displayed on the display device 108. For this purpose, the embodiments aim to reduce the latency from motion to photon, which is the time between when the user moves their head and when the image (photon) incorporating this motion arrives on the user's retina. Without time warping, the latency from motion to photon is the time between the time when the motion that the user generated in pose 122 and the time when the photon is emitted towards the eye 110. With time warping, the latency from motion to photon is the time between the time when the motion that the user generated in the latest pose 124 and the time when the photon is emitted towards the eye 110. In an attempt to reduce the error caused by the latency from motion to photon, the pose estimator may predict the user's pose. As the pose estimator makes a further temporal prediction of the user's pose, also known as the prediction range, the prediction becomes more uncertain. Conventional systems that do not implement time warping in the manner disclosed herein have conventionally had a latency from motion to photon of at least one frame duration or more (e.g., at least 16 milliseconds or more for 60 Hz). The embodiments achieve a latency from motion to photon of about 1 to 2 milliseconds.
[0033] The embodiments disclosed herein are directed to two non - mutually exclusive types of time warping, namely, continuous time warping (CTW) and sequential binocular time warping (SBTW). The embodiments may use a scanning fiber or any other scanning image source (e.g., a micro - electro - mechanical system (MEMS) mirror) as the image source and be used in combination with a display device (e.g., display device 62 of FIG. 1). The scanning fiber relays light from a remote source to the scanning fiber via a single - mode optical fiber. The display device uses an active optical fiber cable to scan - output an image that is much larger than the aperture of the fiber itself. The scanning fiber approach is not bounded by a scanning input start time and a scanning output start time, and a conversion can exist between the scanning input start time and the scanning output start time (e.g., before the image can be uploaded to the display device). Instead, continuous time warping can be implemented, and the conversion is done on a pixel - by - pixel basis, and the x - y location of the pixel or even the x - y - z location is adjusted as the image is slid by the eye.
[0034] The various embodiments discussed herein may be implemented using the system 80 illustrated in FIG. 2. However, the embodiments are not limited to system 80 and may be used in combination with any system capable of implementing the time warping methods discussed herein. (Viewpoint adjustment and warping)
[0035] According to various embodiments, an AR system (e.g., system 80) may use a two-dimensional (2D) see-through display (e.g., display device 62). To represent a three-dimensional (3D) object on the display, the 3D object may need to be projected onto one or more planes. The resulting image on the display may depend on the viewpoint of a user (e.g., user 60) of system 80 looking through the display device 62 at the 3D object. FIGS. 5-7 illustrate the view projection by showing the movement of user 60 relative to the 3D object and what is visible to user 60 at each position.
[0036] FIG. 5 illustrates a user's viewing area from a first position. When the user is positioned at the first position 314, the user can see a first 3D object 304 and a second 3D object 306, as illustrated in "What the user sees 316". From the first position 314, the user can see the first 3D object 304 in its entirety, and a portion of the second 3D object 306 is hidden by the first 3D object 304 placed in front of the second 3D object 306.
[0037] FIG. 6 illustrates the user's viewing area from the second position. When the user translates parallel (e.g., moves laterally) with respect to the first position 314 of FIG. 5, the user's viewpoint changes. Therefore, the features of the first 3D object 304 and the second 3D object 306 visible from the second position 320 may differ from the features of the first 3D object 304 and the second 3D object 306 visible from the first position 314. In the embodiment illustrated in FIG. 6, when the user translates laterally parallel from the second 3D object 306 towards the first 3D object 304, the first 3D object 304 appears to the user to hide a larger portion of the second 3D object 304 as compared to the view from the first position 314. When the user is positioned at the second position 320, the first 3D object 304 and the second 3D object 306 are visible to the user as illustrated in "What the user sees 318". According to various embodiments, when the user translates laterally parallel in the manner illustrated in FIG. 6, "What the user sees 318" is updated non-uniformly (i.e., objects closer to the user (e.g., the first 3D object 304) appear to move more than objects further away (e.g., the second 3D object 306)).
[0038] FIG. 7 illustrates the user's viewing area from the third position. When the user rotates with respect to the first position 314 in FIG. 5, the user's viewpoint changes. Therefore, the features of the first 3D object 304 and the second 3D object 306 visible from the third position 324 may be different from the features of the first 3D object 304 and the second 3D object 306 visible from the first position 314. In the embodiment illustrated in FIG. 7, when the user rotates clockwise, the first 3D object 304 and the second 3D object 306 shift to the left compared to the "content visible to the user 316" from the first position 314. When the user is positioned at the third position 324, the user can see the first 3D object 304 and the second 3D object 306 as illustrated in the "content visible to the user 322". According to various embodiments, when the user rotates around the optical center (e.g., around the center of the viewpoint), the projected image "content visible to the user 322" simply translates. The relative arrangement of the pixels does not change. For example, the relative arrangement of the pixels of the "content visible to the user 316" in FIG. 5 is the same as that of the "content visible to the user 322" in FIG. 7.
[0039] Figures 5-7 illustrate how the rendering of visible objects depends on the user's position. Modifying what the user sees (i.e., the rendered view) based on the user's movement affects the quality of the AR experience. In a seamless AR experience, the pixels representing virtual objects should always appear spatially aligned with the physical world (referred to as "Pixels in the World" (PStW)). For example, if a virtual coffee mug can be placed on a real table in an AR experience, the virtual mug should appear fixed on the table as the user looks around (i.e., changes the viewing point). If PStW is not achieved, the virtual mug will drift in space as the user looks around, thereby disrupting the perception of the virtual mug on the table. In this embodiment, the real table is static with respect to the real-world orientation, while the user's viewing point changes through changes in the user's head pose. Therefore, the system 80 may need to estimate the head pose (with respect to world coordinates), align the virtual objects with the real world, and then draw / present the photons of the virtual objects from the correct viewing point.
[0040] The incorporation of the correct viewing pose into the presented image is important for the PStW concept. This incorporation may occur at different points along the rendering pipeline. Typically, the PStW concept may be better achieved when the time between pose estimation and the presented image is short or when the pose prediction over a given prediction range is more accurate, as the presented image will result in the most up-to-date possible. That is, if the pose estimation becomes stale by the time the image generated based on the pose estimation is displayed, the pixels will not be faithful to the world and PStW cannot be achieved.
[0041] The relationship between the pixel positions of the "content visible to the user 316" corresponding to the first position 314 in FIG. 5 and the pixel positions of the "content visible to the user 318" corresponding to the second position 320 (i.e., the translated position) in FIG. 6 or the "content visible to the user 324" corresponding to the third position (i.e., the rotated position) in FIG. 7 can be referred to as image conversion or warping. (Time warping)
[0042] Time warping can refer to a mathematical transformation between 2D images corresponding to different viewpoints (e.g., the position of the user's head). As the position of the user's head changes, time warping may be applied to transform the displayed image to match the new viewpoint without the need to re-render a new image. Thus, changes in the position of the user's head can be quickly taken into account. Time warping can enable the AR system to appear more responsive and immersive as the user moves their head and thereby modifies their viewpoint.
[0043] In an AR system, after a pose (e.g., a first pose) is estimated based on a user's initial position, the user can move and / or change position, thereby changing what the user sees, and a new pose (e.g., a second pose) may be estimated based on the user's changed position. An image rendered based on the first pose needs to be updated based on the second pose for a realistic VR or AR experience, taking into account the user's movement and / or change in position. To quickly account for this change, the AR system may generate a new pose. Time warping may be performed using the second pose to generate a transformed image that takes into account the user's movement and / or change in position. The transformed image is sent to a display device and displayed to the user. As described above, time warping transforms the rendered image and thus time warping works with existing image data. Therefore, time warping cannot be used to render completely new objects. For example, if the rendered image shows a can and the user moves their head far enough so that a penny hidden behind the can is visible to the user, the penny cannot be rendered using time warping because it was not within the view in the rendered image. However, time warping may be used to appropriately render the position of the can based on the new pose (e.g., according to the new user's viewpoint).
[0044] The effectiveness of time warping can depend on: (1) the accuracy of the new head pose (from which time warping is calculated), i.e., the quality of the pose estimator / sensor fusion, if the warping occurs just before the image is displayed; (2) the accuracy of the pose prediction over the prediction range time, i.e., the quality of the pose predictor, if the warping occurs a short time (prediction range) before the image is displayed; and (3) the length of the prediction range time (e.g., the shorter, the better). (I. Time Warping Operation)
[0045] The time warping operation may include late frame time warping and / or asynchronous time warping. Late frame time warping may refer to warping of a rendered image (or a time warped version thereof) as late as possible within a frame generation cycle before the frame including the rendered image (or its time warped version) is presented to a display (e.g., display device 62). The purpose is to minimize projection error (e.g., error in aligning the virtual world with the real world based on the user's viewpoint) by minimizing the time between when the user's pose is estimated and when the rendered image corresponding to the user's pose is visually recognized by the user (e.g., motion / photon / pose estimate-to-photon latency). By using late frame time warping, the motion-to-photon latency (i.e., the time between when the user's pose is estimated and when the rendered image corresponding to the user's pose is visually recognized by the user) can be less than the frame duration. That is, late frame time warping can be performed quickly, thereby providing a seamless AR experience. Late frame time warping may be executed on a graphics processing unit (GPU). Late frame time warping can work well with a simultaneous / flash panel display that displays all pixels of a frame simultaneously. However, late frame time warping cannot work as well with a sequential / scanning display that displays a frame pixel by pixel as the pixels are rendered. Late frame time warping may be parametric warping, non-parametric warping, or asynchronous warping.
[0046] FIG. 8 illustrates a system for performing latency frame time warping according to one embodiment. As shown in FIG. 8, an application processor / GPU 334 (hereinafter referred to as GPU 334) may perform time warping on image data before transmitting the warped image data (e.g., red-green-blue (RGB) data) 336 to a display device 350. In some embodiments, the image data 336 may be compressed image data. The display device may include a stereoscopic display device. In this embodiment, the display device 350 may include a left display 338 and a right display 340. The GPU 334 may transmit the image data 336 to the left display 338 and the right display 340 of the display device 350. The GPU 334 may have the ability to transmit sequential data per depth and may not need to reduce the data to a 2D image. The GPU 334 includes a time warping module 335 for warping the image data 336 before transmission to the display device 350 (e.g., an eyepiece display such as LCOS).
[0047] Both latency frame time warping and asynchronous time warping may be executed on the GPU 334, and the conversion domain (e.g., the portion of the image that will be converted) may include the entire image (e.g., the entire image is warped simultaneously). After the GPU 334 warps the image, the GPU 334 transmits the image to the display device 350 without further modification. Thus, latency frame or asynchronous time warping may be suitable for applications on a display device that includes a synchronous / flash display (i.e., a display that illuminates all pixels at once). For such a display device, the entire frame must be warped by the GPU 334 before the left display 338 and the right display 340 of the display device 350 are turned on (i.e., the warping must be completed). (II. Continuous Time Warping)
[0048] According to some embodiments, the GPU may render image data and output the rendered image data to an external component (e.g., an integrated circuit such as a field programmable gate array (FPGA)). The external component may perform time warping on the rendered image and output the warped image to a display device. In some embodiments, the time warping may be continuous time warping ("CTW"). CTW may include gradually warping the image data in the external hardware until just before the warped image data is transmitted from the external hardware to the display device and the warped image is converted to photons there. An important feature of CTW is that the continuous warping may be performed on a sub-section of the image data as the image data is streamed from the external hardware to the display device.
[0049] The continuous / streaming operation of CTW may be suitable for applications on a display device, including a sequential / scanning display (i.e., a display that outputs lines or pixels over time). For such a display device, the streaming nature of the display device is coordinated with the streaming nature of CTW, resulting in time efficiency.
[0050] Embodiments provide four exemplary continuous time warping methods, namely, read cursor redirect, pixel redirect, buffer resmear, and write cursor redirect. (1. Read Cursor Redirect (RCRD) Method)
[0051] As used herein, a display device pixel may refer to a display element / unit of a physical display device (e.g., a phosphor cell on a CRT screen). As used herein, an image pixel may refer to a unit of a digital representation of a computer-generated image (e.g., a 4-byte integer). As used herein, an image buffer refers to a region of a physical memory storage device that is used to temporarily store image data while the image data is moved from one location (e.g., memory external to the GPU or a hardware module) to another location (e.g., a display device).
[0052] According to various embodiments, the GPU may input rendered image data (e.g., image pixels) into the image buffer using a write cursor that scans the image buffer by advancing through time within the image buffer. The image buffer may output image pixels to the display device using a read cursor that scans the image buffer by advancing through time within the image buffer.
[0053] With respect to a display device, including a sequential / scanning display, the display device pixels may be turned on in a predefined order (e.g., left to right, top to bottom). However, the image pixels displayed on the display device pixels may vary. As each display device pixel becomes ready to be turned on in sequence, the read cursor may advance through the image buffer and select the image pixel that will be projected next. In the absence of CTW, each display device pixel would always correspond to the same image pixel (e.g., the image is viewed without modification).
[0054] In an embodiment implementing read cursor redirect (``RCRD''), the read cursor may be continuously redirected to select image pixels different from the default image pixels. This results in warping of the output image. When using a sequential / scanning display, the output image may be output line by line, and each line may be warped individually such that the final output image perceived by the user is warped in the desired manner. The displacement of the read cursor and the image buffer may be relative to each other. That is, by using RCRD, when the read cursor is redirected, the image data in the image buffer may be offset. The redirect of the read cursor may be equivalent to a translation of the output image.
[0055] FIG. 9 illustrates an RCRD CTW method according to an embodiment. The same redirect vector may be used for all display device pixels. The read cursor may be directed from the image buffer 332 to the selected image pixel 800 and displayed on the corresponding display device pixel 804 of the display device 350. This may be referred to as the default position of the read cursor or simply the default cursor. By using RCRD, the read cursor may be redirected to the selected image pixel 802 of the image buffer 332 and displayed on the corresponding display device pixel 804 of the display device 350. This may be referred to as the redirected position of the read cursor or simply the redirected cursor. As a result of RCRD, the read cursor selects the image pixel 802 of the image buffer 332 (as further described below in relation to FIG. 11) and transmits the selected image pixel 802 to the display device pixel 804 of the display device 350 for display. When the same redirect vector is applied to all display device pixels of the display device 350, the image data in the image buffer 332 is translated 2 columns to the left. The image pixels of the image buffer 332 are warped, and the resulting displayed image 330 is a translated version of the image data in the image buffer 332.
[0056] Figure 10 illustrates a system for performing RCRD according to one embodiment. An external (i.e., external to the GPU and the display device) controller 342 is provided between the GPU 334 and the display device 350. The GPU 334 may generate image data 336 and transmit it to the external controller 342 for further processing. The external controller 342 may also receive inertial measurement unit (IMU) data 344 from one or more IMUs. In some embodiments, the IMU data 344 may include viewer position data. The external controller 342 may decompress the image data 336 received from the GPU 334, apply continuous-time warping to the decompressed image data based on the IMU data 344, perform pixel conversion and data segmentation, and recompress the resulting data for transmission to the display device 350. The external controller 342 may transmit image data 346 to the left display 338 and image data 348 to the right display 340. In some embodiments, the image data 346 and the image data 348 may be compressed and warped image data.
[0057] In some embodiments, both the left-rendered image data and the right-rendered image data may be transmitted to the respective left display 338 and right display 340. Thus, the left display 338 and the right display 340 may use the additional image data to perform additional accurate image rendering operations such as disocclusion. For example, the right display 340 may perform disocclusion using the image data rendered on the left in addition to the image data to be rendered on the right prior to rendering the image on the right display 340. Similarly, the left display 338 may perform disocclusion using the image data rendered on the right in addition to the image data to be rendered on the left prior to rendering the image on the left display 340.
[0058] FIG. 11 illustrates an external controller 342 as an external hardware unit (e.g., a field programmable gate array (FPGA), a digital signal processor (DSP), an application specific integrated circuit (ASIC), etc.) between a GPU 334 and a display device 350 within a system architecture that performs RCRD CTW. The pose estimator / predictor module 354 of the external controller 342 receives optical data 352 and IMU data 344 from one or more IMUs 345. The external controller 342 may receive image data 336 from the GPU 334 and decompress the image data 336. The decompressed image data may be provided to a (sub) frame buffer 356 of the external controller 342. The read cursor redirect module 396 performs RCRD continuous time warping and converts the compressed image data 336 of the GPU 334 based on the output 387 of the pose estimator / predictor 354. The generated data 346, 348 are time-warped image data, which is then transmitted to the display device 350 and converted into photons 358 emitted towards the viewer's eye. (i. (sub) frame buffer size)
[0059] Referring now to FIGS. 12 and 13, RCRD is discussed with respect to the size of a (sub) frame buffer (e.g., (sub) frame buffer 356). According to some embodiments, the GPU 334 may produce a raster image (e.g., an image as a dot matrix data structure), and the display device 350 may output the raster image (e.g., the display device 350 may "raster output"). As illustrated in FIG. 12, a read cursor 360 (a cursor that advances through the image buffer and selects the image pixels that will be projected next) may advance in a raster pattern through the (sub) frame buffer 356 without RCRD.
[0060] By using an RCRD as illustrated in FIG. 13, the re-estimation of the viewer's pose immediately before an image pixel is displayed can incorporate taking the image pixel to be presented to the user from a different position within the buffered image instead of taking the image pixel from the default position. Assuming bounded head / eye movement, the locus of the redirected reading position (i.e., the different position) is a closed set within a bounded area (e.g., circular B) 362 around the reading cursor 360. This locus is superimposed with raster forward momentum because the display device 350 is still raster outputting the image. FIG. 13 shows that by using continuous warping / pose re-estimation, the reading cursor trajectory is the superposition of the locus and raster forward movement. The buffer height 364 of the (sub) frame buffer 356 should be equal to or exceed the diameter of the bounded area 362 to prevent the reading cursor 360 from extending beyond the boundaries of the (sub) frame buffer 356. A larger (sub) frame buffer may require additional processing time and computational power.
[0061] The diameter of the bounded area 362 is a function of the rendering rate of the display device 350, the accuracy of the pose prediction at the rendering time, and the maximum speed of the user's head movement. Thus, for a higher rendering rate, since the higher rendering rate results in a shorter elapsed time for the head pose to deviate from the pose assumption at the rendering time, the diameter of the bounded area 362 decreases. Accordingly, the redirect that may be required for the reading cursor may be small (e.g., the redirected reading position will be closer to the original reading cursor position). Additionally, for a more accurate pose prediction at the rendering time (e.g., when an image is being rendered), since the required time warping correction will be smaller, the diameter of the bounded area 362 decreases. Accordingly, the redirect that may be required for the reading cursor may be small (e.g., the redirected reading position will be closer to the original reading cursor position). Further, for a higher head movement speed, since the head pose may deviate more over a given time interval with high head movement, the diameter of the bounded area 362 increases. Accordingly, the redirect that may be required for the reading cursor may be substantial (e.g., the redirected reading position will be farther from the original reading cursor position). (ii. Reading cursor vs. writing cursor position)
[0062] According to some embodiments, the (sub) frame buffer 356 may also include a write cursor. In the case without read cursor redirect, the read cursor 360 follows immediately after the write cursor and can move, for example, both in the raster forward direction through the (sub) frame buffer 356. The (sub) frame buffer 356 may include a first area (e.g., a new data area) for the content rendered at a certain timestamp and a second area (e.g., an old data area) for the content rendered at the previous timestamp. FIG. 14 illustrates the (sub) frame buffer 356 including a new data area 368 and an old data area 370. By using the RCRD, when the read cursor 360 reading from the (sub) frame buffer 356 follows immediately after the write cursor 366 writing to the (sub) frame buffer 356, the trajectory of the redirected read position (e.g., as depicted by the bounded area 362 in FIG. 13) can result in an "area intersection" where the read cursor 360 intersects into the old data area 370 as illustrated in FIG. 14. That is, the read cursor 360 may read the old data in the old data area 370 and may not be able to read the new data being written to the (sub) frame buffer 356 by the write cursor 366.
[0063] Region intersections can result in image tearing. For example, if the image pixels in the (sub) frame buffer 356 include the depiction of a straight vertical line moving to the right, and the read cursor 360 moves back and forth between two content renderings (e.g., the first in the new data area 368 and the second in the old data 370) due to RCRD, the line that is displayed will not be straight and will have a tear if a region intersection occurs. Image tearing can be prevented by centering the read cursor 360 behind the write cursor 366 such that the bounded area 362 at the boundary of the redirected position of the read cursor 360 is within the new data area 368 as shown in FIG. 15. This can be accomplished by setting the buffer read distance 372 between the center of the bounded area 362 and the boundary 639 separating the new data area 368 from the old data area 370. The buffer read distance 372 can keep the bounded area 362 completely within the new data area 368.
[0064] Repositioning of the read cursor 360 achieves the desired output image pixel orientation (i.e., the orientation of the image pixels is accurate regardless of content timing). Thus, positioning the bounded area 362 behind the write cursor 366 does not adversely affect the pose estimate / prediction-to-photon latency from PStW or pose estimation / prediction.
[0065] On the other hand, the buffer read distance 372 can increase the rendering-to-photon latency in proportion to the buffer read distance 372. The rendering-to-photon latency is the time between the scene rendering time and the photon output time. Photons can have a latency from zero pose estimation / prediction to the photon by perfect time warping, but the rendering-to-photon latency can only be reduced by reducing the time between scene rendering and photon output time. For example, for a 60 frames per second (fps) rendering rate, the default rendering-to-photon latency can be about 16 milliseconds (ms). 10 lines of buffer read (e.g., buffer read distance 372) in a 1,000-line image can only add 0.16 ms of rendering-to-photon latency. According to some embodiments, the rendering-to-photon latency increase is removed when buffer transmission time is not required (e.g., when there is no write cursor), for example, when the external control 342 (e.g., FPGA) directly accesses the GPU 334, thereby eliminating the rendering-to-photon latency increase caused by the transmission time between the GPU 334 and the external control 342. (iii. External Anti-Aliasing)
[0066] Anti-aliasing is a view-dependent operation that is performed after time warping to avoid blur artifacts. By using CTW, anti-aliasing can be performed by external hardware immediately before photon generation. External control 342 (e.g., FPGA) with direct access to the GPU 334 (or two cooperating GPUs) may be used to perform external anti-aliasing after continuous time warping and before photon generation for display. (2. Buffer Resmear Method)
[0067] RCRD CTW may not need to consider the translational movement of the user's head, which requires non-uniformly shifting image pixels according to the pixel depth (or the distance to the viewer). Different CTW methods, namely, the buffer resmear method, may be used to render the image even when the viewer is translating. Buffer resmear is a concept that incorporates the latest pose estimation / prediction by updating the buffered image pixels before extracting the image pixels for the read cursor to be displayed. Since different image pixels can be shifted by different amounts, buffer resmear can take into account the translational movement of the user's head. FIG. 16 illustrates a buffer resmear CTW method according to one embodiment. Buffer resmear using the latest pose implemented on the image buffer 374 results in a modified image buffer 376. Image pixels from the modified image buffer 376 are displayed on the corresponding display device pixels of the display device 350.
[0068] FIG. 17 illustrates a system architecture for buffer resmear CTW according to one embodiment. An external controller 342 is provided between the GPU 334 and the display device 350. The pose estimator / predictor module 354 of the external controller 342 receives the optical data 352 and the IMU data 344 (from one or more IMUs 345). The external controller 342 receives the compressed image data 336 from the GPU 334 and decompresses the image data 336. The decompressed image data may be provided to the (sub) frame buffer 356 of the external controller 342. The external buffer processor 378 performs buffer resmear on the compressed image data 336 received from the GPU 334 based on the output 387 of the pose estimator / predictor 354 before the pixels are transmitted to the display device 350 and converted into photons 358 that are emitted towards the viewer's eyes.
[0069] According to various embodiments, buffer resampling can occur each time a new pose is to be incorporated into the displayed image, which can be on a per-pixel basis for a sequential display. Even if only a portion of the (sub)frame buffer 356 is resampled, this is a computationally costly operation. (3. Pixel Redirection Method)
[0070] Another CTW method can be a pixel redirection method, which is the reverse operation of the read cursor redirection method. According to the pixel redirection method, instead of the display device 350 determining the appropriate image pixels to fetch, the external controller determines the display device pixels that are to be activated for a given image pixel. In other words, the external controller determines the display device pixels where the image pixel needs to be displayed. Thus, in pixel redirection, each image pixel can be repositioned independently. FIG. 18 illustrates that pixel redirection can result in warping that also takes into account the rotation and / or (partially) translational movement of the user's head as well.
[0071] As shown in FIG. 18, the first image pixel 391 in the image buffer 380 may originally be defined to be displayed on the first display device pixel 393 of the display device 350. That is, the first display device pixel 393 may be assigned to the first image pixel 391. However, the pixel redirect method may determine that the first image pixel 391 should be displayed on the second display device pixel 395, and the external controller may send the first image pixel 391 to the second display device pixel 395. Similarly, the second image pixel 394 in the image buffer 380 may originally be defined to be displayed on the third display device pixel 397 of the display device 350. That is, the second image pixel 394 may be assigned to the third display device pixel 397. However, the pixel redirect method may determine that the second image pixel 394 should be displayed on the fourth display device pixel 399, and the external controller may send the second image pixel 394 to the fourth display device pixel 399. Pixel redirect implemented on the image buffer 380 results in the resulting displayed image 382 being displayed on the display device 350. Pixel redirect may require a special type of display device 350 that can selectively turn on arbitrary pixels in an arbitrary order. A special type of OLED or similar display device may be used as the display device 350. The pixels may first be redirected to a second image buffer, and then the second buffer may be sent to the display device 350.
[0072] FIG. 19 illustrates a system architecture for an external hardware pixel redirect method according to one embodiment. An external controller 342 is provided between a GPU 334 and a display device 350. An attitude estimator / predictor module 354 of the external controller 342 receives optical data 352 and IMU data 344 (from one or more IMUs 345). The external controller 342 may receive image data 336 from the GPU 334 and decompress the image data 336. The decompressed image data may be provided to a (sub) frame buffer 356 of the external controller 342. The output 387 of the attitude estimator / predictor 354 and the output 389 of the (sub) frame buffer 356 are provided to the display device 350 and converted into photons 358 emitted towards the viewer's eyes. (4. Writing Cursor Redirect Method)
[0073] Another CTW method, namely, Write Cursor Redirect (WCRD) method, changes the way image data is written into the (sub) frame buffer. FIG. 20 illustrates the WCRD method which can also take into account the rotation and (partially) translation of the user's head. Each pixel can be repositioned independently (e.g., using forward mapping / scattering operations). For example, the first image pixel 401 in the image buffer 333 of the GPU 334 could originally be defined as the first image pixel 403 in the (sub) frame buffer 356 of an external controller 342 (e.g., FPGA). However, by using forward mapping, the first image pixel 401 can be directed to the second image pixel 404 in the (sub) frame buffer 356. Similarly, the second image pixel 402 in the image buffer 333 could originally be defined as the third image pixel 405 in the (sub) frame buffer 356. However, by using forward mapping, the second image pixel 402 can be directed to the fourth image pixel 406 in the (sub) frame buffer 356. Thus, the image can be warped during the data transfer from the frame buffer 333 of the GPU 334 to the (sub) frame buffer 356 of the external controller 342 (e.g., FPGA). That is, CTW is performed on the image before the image reaches the (sub) frame buffer 356.
[0074] Figure 21 illustrates a system architecture for an external hardware WCRD CTW according to one embodiment. An external controller 342 is provided between a GPU 334 and a display device 350. The pose estimator / predictor module 354 of the external controller 342 receives optical data 352 and IMU data 344 (from one or more IMUs 345). Image data 336 is transmitted by the GPU 334 (i.e., the frame buffer 333 of the GPU 334), and the output 387 of the pose estimator / predictor 354 is received at the write cursor redirect module 386 of the external controller 342. For each incoming image data pixel, the image pixel is redirected to and written at a pose-matching location within the (sub) frame buffer 356 based on the current pose estimation / prediction and the depth of that image pixel. The outputs 346, 348 of the (sub) frame buffer 356 are time-warped image data, which is then transmitted to the display device 350 and converted into photons 358 emitted towards the viewer's eye.
[0075] According to various embodiments, the write cursor redirect module 386 may be a 1-pixel buffer since the external controller 342 needs to process the location where the image pixel is to be written.
[0076] Figure 22 illustrates the WCRD method, where the write cursor 366 has a locus of write positions that is a closed set within a bounded area (e.g., circular B) 388 that advances through the (sub) frame buffer 356. The buffer height 392 of the (sub) frame buffer 356 needs to be at least as large as the diameter of the bounded area 388. The WCRD method may require a buffer read distance 390 between the center of the bounded area 388 and the read cursor 360. According to various embodiments, the buffer read distance 390 may be a function of at least one of the frame rate of the display device, the resolution of the image, and the expected speed of head movement.
[0077] According to some embodiments, the WCRD method can introduce some latency from pose estimation / prediction to photons because pose estimation / prediction can incorporate an amount of time (proportional to the buffer read distance) before photons are generated. For example, for a 60fps display clock output rate, a 10-line buffer of a 1,000-line image can introduce a latency of 0.16ms from pose estimation / prediction to photons. (5. Write / Read Cursor Redirection Method)
[0078] The CTW methods discussed herein may not be mutually exclusive. By complementing the WCRD method with the RCRD method, the rotation and translation of the viewer can be taken into account. FIG. 23 illustrates that by using the write-read cursor redirect (WRCRD) method, both the write and read cursor positions are within the bounded areas 388 and 362, respectively, but the bounded area 362 of the read cursor 360 is much smaller compared to the bounded area 388 of the write cursor 366. The minimum buffer height 392 can be determined to accommodate both the bounded area 362 and the bounded area 388 without causing an intersection from the new data area 368 to the old data area 370. In some embodiments, the minimum buffer height 392 for the WRCRD method may be twice the buffer height 364 for the RCRD method. Additionally, the buffer read distance 410 can also be determined to accommodate both the bounded area 362 and the bounded area 388 without causing an intersection from the new data area 368 to the old data area 370. In some embodiments, the buffer read distance 410 for the WRCRD method can be twice (e.g., 20 lines) the buffer read distance 390 (e.g., 10 lines) for WCRD or the buffer read distance 372 (e.g., 10 lines) for RCRD.
[0079] According to some embodiments, the size of the bounded area 388 is proportional to the amount of pose adjustment required after the last pose incorporation during rendering. The size of the bounded area 362 is also proportional to the amount of pose adjustment required after the last pose incorporation when writing pixel data to the (sub) frame buffer. When the buffer distance of the read cursor to the write cursor is 10 lines within a 1,000-line image, the elapsed time between the time when image data is written into a pixel by the write cursor 366 and the time when the image data is read from the pixel by the read cursor 360 is about 1% of the elapsed time between the pose estimation of the write cursor 366 and the pose estimation at the rendering time. In other words, when the read cursor 360 is closer to the write cursor 366, the read cursor 360 reads more recent data (e.g., data written more recently by the write cursor 366), thereby shortening the time between when the image data is written and when the image data is read. Therefore, the buffer size and the read distance do not need to be doubled and can only be increased by a few percent.
[0080] RCRD in the WRCRD method cannot consider the translational movement of the user's head. However, WCRD in the somewhat earlier-occurring WRCRD method considers the translational movement of the user's head. Therefore, WRCRD can achieve very short (e.g., virtually zero) latency parameter warping and very short latency non-parameter warping (e.g., about 0.16 ms for the display clock output at 60 fps and a 10-line buffer of a 1,000-line image).
[0081] The system architecture for an external hardware WRCRD CTW according to one embodiment is illustrated in FIG. 24. An external controller 342 is provided between a GPU 334 and a display device 350. The pose estimator / predictor module 354 of the external controller 342 receives optical data 352 and IMU data 344 (from one or more IMUs 345). The image data 336 transmitted by the GPU 334 (i.e., the frame buffer 333 of the GPU 334) and the output 387 of the pose estimator / predictor 354 are received in a write cursor redirect module 386. With respect to incoming image data, each image pixel is redirected to and written at a pose-matching location within a (sub) frame buffer 356 based on the current pose estimation / prediction and the depth of that image pixel. Additionally, a read cursor redirect module 396 performs WRCRD based on the output 387 of the pose estimator / predictor 354 and converts the image data received from the write cursor redirect module 386. The generated data 346, 348 are time-warped image data, which are then transmitted to the display device 350 and converted into photons 358 emitted towards the viewer's eye. According to various embodiments, the WRCRD method can be implemented on the same external controller 342, act on a single (sub) frame buffer 356, and act on streamed data. The write cursor redirect module 386 and / or the read cursor redirect module 396 may be independently turned off for different display options. Thus, the WRCRD architecture can function as an on-demand WCRD or RCRD architecture. (III. Binocular Time Warping)
[0082] As used herein, binocular time warping refers to latency frame time warping, which is used in combination with a display device including a left display unit for the left eye and a right display unit for the right eye, and the latency frame time warping is performed separately for the left display unit and the right display unit. FIG. 25 illustrates binocular time warping. Once 3D content is rendered in the GPU 3002 at time t1 3003 and the latest pose input 3004 is received from the pose estimator 3006 before time t2 3007, time warping is performed for both the left frame 3008 and the right frame 3010 at time t2 3007 simultaneously or almost simultaneously. For example, in an embodiment where both time warping is performed by the same external controller, the time warping for the left frame and the time warping for the right frame may be performed consecutively (e.g., almost simultaneously).
[0083] The transformed images 3014 and 3016 (i.e., the images for which time warping is performed) are transmitted to the left display unit 3018 and the right display unit 3020 of the display device 350, respectively. Photons are generated in the left display unit 3018 and the right display unit 3020 and emitted towards the individual eyes of the viewer, thereby displaying the images simultaneously (e.g., at time t3 3015) on the left display unit 3018 and the right display unit 3020. That is, in one embodiment of binocular time warping, the same latest pose 3004 is used to perform time warping on the same rendered frame for both the left display unit 3018 and the right display unit 3020.
[0084] In another embodiment, alternating binocular time warping may be used to perform time warping for different latest poses for the left display unit 3018 and the right display unit 3020. The alternating time warping may be performed in various manners, as illustrated in FIGS. 26 to 29. The alternating binocular time warping illustrated in FIGS. 26 to 29 enables presenting the latest pose update information regarding the image displayed on the display device 350. An old frame (i.e., a previously displayed frame or a frame received from the GPU) may be used to interpolate the time warping. By using the alternating binocular time warping, the latest pose can be incorporated into the displayed image, and the latency from motion to photons can be reduced. That is, instead of both being updated simultaneously using a "more recent" latest pose, one eye can view the image using a pose that is incorporated into the warping later than the other.
[0085] FIG. 26 illustrates another alternating binocular time warping according to an embodiment. The same GPU 3070 is used to generate both the left and right rendered viewpoints, which are respectively used by the left display unit 3018 and the right display unit 3020 at time t1 3071. The first time warping is performed on the rendered left frame by the time warping left frame module 3072 at time t2 3073 using the first latest pose 3074 received from the pose estimator 3006. The output of the time warping left frame module 3072 is transmitted to the left display unit 3018. The left display unit 3018 converts the received data into photons at time t4 3077 and emits the photons towards the viewer's left eye, thereby displaying the image on the left display unit 3018.
[0086] The second time warping is performed on the rendered right frame by the time warping right frame module 3078 at time t3 3079 (e.g., a time after t2 3073 at which the first time warping is performed) using the second latest pose 3080 received from the pose estimator 3006. The output of the time warping right frame module 3078 is transmitted to the right display unit 3020. The right display unit 3020 converts the received data into photons at time t5 3081 and emits the photons towards the viewer's right eye, thereby displaying an image on the right display unit 3020. The right display unit 3020 displays the image after the time when the left display unit 3018 displays the image.
[0087] FIG. 27 illustrates another type of alternating binocular time warping according to an embodiment. The same GPU 3050 is used to generate the rendered frames for both the left display unit 3018 and the right display unit 3020. The left frame and the right frame are rendered at different times (i.e., the left frame and the right frame are rendered alternately in time). As shown, the left frame may be rendered at time t1 3051, and the right frame may be rendered at time t2 3052 after time t1 3051. The first time warping is performed on the rendered left frame by the time warping left frame module 3058 at time t3 3053 using the first latest pose 3054 received from the pose estimator 3006 before time t3 3053. The output of the time warping left frame module 3058 is transmitted to the left display unit 3018. The left display unit 3018 converts the received data into photons at time t5 3061 and emits the photons towards the viewer's left eye, thereby displaying an image on the left display unit 3018.
[0088] The second time warping is performed on the rendered right frame by the time warping right frame module 3060 at time t4 3059 (e.g., after the time t3 3053 at which the first time warping is performed) using the second latest pose 3062 received from the pose estimator 3006 before time t4 3059. The output of the time warping right frame module 3060 is transmitted to the right display unit 3020. The right display unit 3020 converts the received data into photons at time t6 3063 and emits the photons towards the viewer's right eye, thereby displaying an image on the right display unit 3020. The right display unit 3020 displays the image at a time after the left display unit 3018 displays the image.
[0089] FIG. 28 illustrates alternating binocular time warping according to one embodiment. According to the embodiment illustrated in FIG. 28, two separate GPUs 3022 and 3024 may be used to generate the rendered views for the left display unit 3018 and the right display unit 3020. The first GPU 3022 may render the left view at time t1 3025. The second GPU 3024 may render the right view at time t2 3026 after time t1 3025. The first time warping is performed on the rendered left view by the time warping left frame module 3030 at time t3 3027 using the first latest pose 3004 received from the pose estimator 3006 before time t3 3027. The output of the time warping left frame module 3030 is transmitted to the left display unit 3018. The left display unit 3018 converts the received data into photons and emits the photons towards the viewer's left eye. The left display unit 3018 displays the image at time t5 3033.
[0090] The second time warping is performed on the rendered right view by the time warping right frame module 3032 at time t4 3031 (e.g., a time after the first time warping is performed), using the second latest pose 3034 received from the pose estimator 3006 before time t4 3031 (e.g., a time after the first latest pose 3004 is obtained by the time warping left frame module 3030). The output of the time warping right frame module 3032 is transmitted to the right display unit 3020. The right display unit 3020 converts the received data into photons and emits the photons towards the viewer's right eye. The right display unit 3020 displays an image at time t6 3035 (i.e., a time after the left display unit 3018 displays an image at time t5 3033). The image displayed on the right display unit 3020 may be more up-to-date because it is generated considering the more recent pose (i.e., the second latest pose 3034).
[0091] FIG. 29 illustrates another type of binocular time warping according to an embodiment. The same GPU 3036 may be used to generate the rendered views for both the left display unit 3018 and the right display unit 3020. As illustrated in FIG. 29, the GPU 3036 may generate a rendered view of the left frame at time t1 3037. The time warped image of the left frame may be generated at time t3 3041. The display update rate of the left display unit 3018 may be slow enough such that the second rendering (i.e., the rendering for the right display unit 3020) can be performed by the GPU 3036 after the time warped image of the left frame is generated (e.g., the time warped image of the left frame is displayed on the left display unit 3018).
[0092] The first time warping is performed on the rendered left frame at time t2 3038 by the time warping left frame module 3040 using the first latest pose 3039 received from the IMU. The output of the time warping left frame module 3040 is transmitted to the left display unit 3018. The left display unit 3018 converts the received data into photons at time t3 3041 and emits the photons towards the viewer's left eye, thereby displaying the image on the left display unit 3018. After the image is displayed on the left display unit 3018, the GPU 3036 may render the right frame at time t4 3042 using the obtained data (e.g., the image and the data received from the IMU for generating the pose estimation regarding the right eye and the 3D content generated from the pose estimation). The second time warping is performed on the rendered right frame at time t5 3043 (e.g., after the time warped image is displayed on the left display unit 3018 at time t3 3041) by the time warping right module 3055 using the second latest pose 3044 received from the IMU. The output of the time warping right frame module 3046 is transmitted to the right display unit 3020. The right display unit 3020 converts the received data into photons at time t6 3047 and emits the photons towards the viewer's right eye, thereby displaying the image on the right display unit 3020. The right display unit 3020 displays the image "x" seconds after the data is received from the image and the IMU, where x is a mathematical relationship less than the refresh rate required for smooth viewing. Therefore, the two display units (i.e., the left display unit 3018 and the right display unit 3020) are updated completely offset from each other, and each update is sequential with respect to the other (e.g., 2x < the refresh rate required for smooth viewing).
[0093] Those skilled in the art will understand that the order in which the left display unit and the right display unit display images may be different from that described above with respect to FIGS. 26 to 29. The system may be modified such that the left display unit 3018 displays an image after the time when the right display unit 3020 displays an image.
[0094] Also, the examples and embodiments described in this specification are for illustrative purposes only, and in light of this, it should be understood that various modifications or changes may be suggested to those skilled in the art and should be included within the spirit and scope of this application and the scope of the appended claims.
[0095] Also, the examples and embodiments described in this specification are for illustrative purposes only, and in light of this, it should be understood that various modifications or changes may be suggested to those skilled in the art and should be included within the spirit and scope of this application and the scope of the appended claims.
Claims
1. 1. A method for time warping rendered image data based on an updated position of a viewer, the method comprising: obtaining, from a graphics processing unit, rendered image data corresponding to a first viewpoint associated with a first pose estimated based on a first position of the viewer; receiving, by the computing device, data associated with a second position of the viewer; and the computing device estimating a second pose based on the second position of the viewer; and the computing device time warping, in a first progression, the first set of rendered image data into a first set of time-warped rendered image data, the time-warped rendered image data corresponding to a second viewpoint associated with the second pose estimated based on the second position of the viewer; the computing device time warping the second set of rendered image data into the second set of time warped rendered image data in a second progression, the first set of time warped rendered image data being different from the second set of time warped rendered image data; transmitting, by the computing device, at a first time, the first set of the time warped rendered image data to a display module of an eyepiece display device; transmitting, by the computing device, at a second time, the second set of time-warped rendered image data to the display module of the eyepiece display device; Including, wherein time warping the second set of rendered image data into the second set of time warped rendered image data is progressive to the first time and is completed by the second time.
2. The method of claim 1 , wherein the second position corresponds to a rotation about an optical center of the eyepiece display device from the first position.
3. The method of claim 1 , wherein the second position corresponds to a horizontal translation from the first position.
4. The method of claim 1 , wherein the rendered image data is progressively time warped into the time warped rendered image data until the time warped rendered image data is converted to photons.
5. The method of claim 1 , wherein the rendered image data is a subsection of an image streamed from the graphics processing unit.
6. The computing device includes a frame buffer for receiving the rendered image data from the graphics processing unit, and the method further comprises: redirecting display device pixels of the eyepiece display device from default image pixels in the frame buffer to different image pixels in the frame buffer, The method of claim 1 further comprising:
7. The computing device includes a frame buffer for receiving the rendered image data from the graphics processing unit, and the method further comprises: the computing device transmitting image pixels in the frame buffer to display device pixels of the eyepiece display device that differ from default display device pixels assigned to the image pixels. The method of claim 1 further comprising:
8. The computing device includes a frame buffer for receiving the rendered image data from the graphics processing unit, and the method further comprises: receiving the rendered image data from a frame buffer of the graphics processing unit while time warping the first set of rendered image data into the first set of time warped rendered image data; Further comprising: The time warping may include: the computing device receiving from the graphics processing unit a first image pixel in a frame buffer of the graphics processing unit at a first image pixel in a frame buffer of the computing device, the first image pixel in the frame buffer of the graphics processing unit being initially assigned to a second image pixel in the frame buffer of the computing device; The method of claim 1 , comprising:
9. 2. The method of claim 1, wherein the data associated with the second position of the viewer includes optical data and data from an inertial measurement unit, and the computing device includes a pose estimator module for receiving the optical data and the data from the inertial measurement unit and estimating the second pose based on the second position of the viewer, and a frame buffer for receiving the rendered image data from the graphics processing unit.
10. 2. The method of claim 1, wherein the computing device includes a frame buffer for receiving the rendered image data from the graphics processing unit, and an external buffer processor for time warping at least the first set of the rendered image data into the first set of the time warped rendered image data by shifting buffered image pixels stored in the frame buffer prior to transmitting the time warped rendered image data to the eyepiece display device.
11. 11. The method of claim 10, wherein a first image pixel in the frame buffer is shifted by a first amount and a second image pixel in the frame buffer is shifted by a second amount different from the first amount.
12. 1. A system comprising: a graphics processing unit configured to generate rendered image data corresponding to a first viewpoint associated with a first pose estimated based on a first position of a viewer; and A controller, the controller comprising: receiving data associated with a second position of the viewer; and estimating a second pose based on the second position of the viewer; and in a first progression, time warping the first set of rendered image data into a first set of time warped rendered image data, the time warped rendered image data corresponding to a second viewpoint associated with the second pose estimated based on the second position of the viewer; in a second progression, time warping the second set of rendered image data into the second set of time warped rendered image data, the first set of time warped rendered image data being different from the second set of time warped rendered image data; transmitting the first set of time warped rendered image data to a display module of an eyepiece display device at a first time; transmitting the second set of time warped rendered image data to the display module of the eyepiece display device at a second time; wherein time warping the second set of rendered image data into the second set of time warped rendered image data is progressive to the first time and is completed by the second time; 1. An eyepiece display device, comprising: receiving the first set of time warped rendered image data and the second set of time warped rendered image data transmitted by the controller; converting the time warped rendered image data to photons; emitting the photons towards the viewer and displaying the time-warped rendered image data on the eyepiece display device; and an eyepiece display device configured to A system comprising:
13. The system of claim 12 , wherein the eyepiece display device is a sequential scan display device.
14. The system of claim 12 , wherein the controller is integrated within the eyepiece display device.
15. The system of claim 12 , wherein the controller is coupled to the graphics processing unit and the eyepiece display device and is provided between the graphics processing unit and the eyepiece display device.
16. The controller: a frame buffer for receiving the rendered image data from the graphics processing unit; an external buffer processor for time warping at least the first set of the rendered image data into the first set of time warped rendered image data by shifting buffered image pixels stored in the frame buffer prior to transmitting the time warped rendered image data to the eyepiece display device; The system of claim 12 further comprising:
17. The controller:
17. The system of claim 16, further configured to redirect display device pixels of the eyepiece display device from default image pixels in the frame buffer to different image pixels in the frame buffer.
18. The controller: The system of claim 16 , further configured to transmit image pixels in the frame buffer to display device pixels of the eyepiece display device that differ from default display device pixels assigned to the image pixels.
19. The system of claim 12 , wherein the second position corresponds to a rotation about an optical center of the eyepiece display device from the first position or a horizontal translation from the first position.
20. The system comprises: An optical unit; Inertial Measurement Unit and Further equipped with The controller: Further comprising a pose estimator module, The pose estimator module: receiving optical data from the optical unit and data from the inertial measurement unit; estimating the second pose based on the second position of the viewer; and The system of claim 12 configured to:
21. The computing device, comprising: a first time warping frame module and a second time warping frame module; The method comprises: the first time warping frame module transmitting the first set of time warped rendered image data to the display module of the eyepiece display device at the first time; the second time warping frame module transmitting the second set of time warped rendered image data to the display module of the eyepiece display device at the second time. The method of claim 1 , comprising:
22. The method of claim 20, wherein the controller comprises a first time warping frame module and a second time warping frame module; the first time warping frame module transmitting the first set of time warped rendered image data to the display module of the eyepiece display device at the first time; the second time warping frame module transmitting the second set of time warped rendered image data to the display module of the eyepiece display device at the second time. The system of claim 12 configured to:
Citation Information
Patent Citations
Program, information storage medium, and image generation system
JP2013062731A
Image generation apparatus and image generation method
JP2015095045A
Image processing
US20160035140A1