Methods for improving the perceptual quality of foveated rendered images
Patent Information
- Application Number
- JP2024508432
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-13
- Filing Date
- 2022-08-12
- Publication Date
- 2025-08-26
AI Technical Summary
Existing foveated rendering technologies face challenges in achieving high accuracy, speed, and low latency for eye tracking, leading to artifacts and reduced viewer experience due to inaccuracies and latencies in eye tracking or content estimation.
Implementing temporal and/or binocular multiplexing of spatial profiles in foveated rendering, which dynamically shifts foveal zones to reduce perceptual artifacts without requiring precise eye tracking, using frame-specific control packets to enhance visual quality and save computational resources.
Enhances perceived visual quality by reducing artifacts and maintaining high resolution across the field of view, while minimizing computational power and bandwidth requirements.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 232,787, entitled “METHODS TO IMPROVE THE PERCEPTUAL QUALITY OF FOVEATED RENDERED IMAGES,” filed on August 13, 2021, the entire contents of which are incorporated herein by reference for all purposes.
[0002] 2. Background of the Invention In the human eye, the fovea is responsible for sharp central vision at the center of gaze. Peripheral vision is vision occurring away from the center of gaze. Visual acuity is lower in peripheral vision compared to foveal vision. Foveated rendering (FR) is a rendering technique in which image resolution, or the amount of detail, is higher in the area of the image corresponding to the fixation point and lower away from the fixation point. FR can achieve a significant reduction in rendering power and bandwidth, which can be advantageous in applications with limited resources, such as virtual reality (VR) and augmented reality (AR).
[0003] Some FR techniques involve tracking the viewer's gaze in real time using an eye-tracking device integrated with the VR / AR headset. For a satisfying viewer experience with FR, the eye-tracking needs to have sufficiently high accuracy, high speed, and low latency, which can be difficult to achieve. Some FR techniques do not use eye-tracking, but instead use a fixed focus. Such FR techniques are called fixed FR. For example, assuming that the viewer looks at the center of the display, the field of view (FOV) of the display can be divided into a central zone with full resolution and several peripheral zones with reduced resolution. The viewer may not always look at the center of the display, which can lead to a compromised viewer experience. Other techniques, such as content-based FR, also do not require eye-tracking, but may require heavy computational resources.
[0004] Thus, there is a need in the art for improved FR techniques. Summary of the Invention [Means for solving the problem]
[0005] Summary of the Invention According to some embodiments, a method of generating a foveated rendering using temporal multiplexing includes generating a first spatial profile of the FOV by dividing the FOV into a first foveal zone and a first peripheral zone. The first foveal zone is rendered at a first pixel resolution and the first peripheral zone is rendered at a second pixel resolution lower than the first pixel resolution. The method further includes generating a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone. The second foveal zone is spatially offset from the first foveal zone. The second foveal zone is rendered at the first pixel resolution and the second peripheral zone is rendered at the second pixel resolution. The method further includes temporally multiplexing the first spatial profile and the second spatial profile with a series of frames such that a viewer perceives an image rendered in a region of the first foveal zone that does not overlap with the second foveal zone and / or a region of the second foveal zone that does not overlap with the first foveal zone when rendered at the first pixel resolution.
[0006] According to some embodiments, a method of generating foveated rendering using binocular multiplexing includes generating a first spatial profile of the FOV by dividing the FOV into a first foveal zone and a first peripheral zone. The first foveal zone is rendered at a first pixel resolution and the first peripheral zone is rendered at a second pixel resolution lower than the first pixel resolution. The method further includes generating a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone. The second foveal zone is spatially offset from the first foveal zone. The second foveal zone is rendered at the first pixel resolution and the second peripheral zone is rendered at the second pixel resolution. The method further includes multiplexing the first spatial profile and the second spatial profile for each of the left and right eyes of the viewer such that, when rendered at the first pixel resolution, the viewer perceives an image rendered in an area of the first foveal zone that does not overlap with the second foveal zone and / or in an area of the second foveal zone that does not overlap with the first foveal zone.
[0007] According to some embodiments, a method of generating foveated rendering using a combination of temporal and binocular multiplexing includes generating a first spatial profile of the FOV by dividing the FOV into a first foveal zone and a first peripheral zone. The first foveal zone is rendered at a first pixel resolution and the first peripheral zone is rendered at a second pixel resolution lower than the first pixel resolution. The method further includes generating a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone. The second foveal zone is spatially offset from the first foveal zone. The second foveal zone is rendered at the first pixel resolution and the second peripheral zone is rendered at the second pixel resolution. The method further includes generating a third spatial profile of the FOV by dividing the FOV into a third foveal zone and a third peripheral zone. The third foveal zone is spatially offset from the first foveal zone. The third foveal zone is rendered at the first pixel resolution and the third peripheral zone is rendered at the second pixel resolution. The method further includes generating a fourth spatial profile of the FOV by dividing the FOV into a fourth foveal zone and a fourth peripheral zone. The fourth foveal zone is spatially offset from the third foveal zone. The fourth foveal zone is rendered at the first pixel resolution and the fourth peripheral zone is rendered at the second pixel resolution. The method further includes multiplexing the first spatial profile for the left eye and the second spatial profile for the right eye of the viewer, respectively, in odd frames, and multiplexing the third spatial profile for the left eye and the fourth spatial profile for the right eye of the viewer, respectively, in even frames.
[0008] According to some embodiments, a method for implementing a video pipeline implementation of dynamically multiplexed foveated rendering includes rendering a foveated image including a foveal zone and a peripheral zone, where the foveal zone has a first set of image data and is rendered at a first pixel resolution, and the peripheral zone has a second set of image data and is rendered at a second pixel resolution lower than the first pixel resolution. The method further includes packing the first set of image data into a first image block and packing the second set of image data into a second image block. The method also includes generating a control packet including rendering information associated with the foveal image, concatenating the control packet with the first image block and the second image block to form a frame, and transmitting the frame to a display unit. The control packet is parsed from the transmitted frame and decoded to obtain the rendering information. Finally, a display image is rendered according to the decoded rendering information.
[0009] According to some embodiments, a method for implementing a video pipeline implementation of dynamic multiplexed foveated rendering with time warping includes rendering a foveated image including a foveal zone and a peripheral zone, where the foveal zone has a first set of image data and is rendered at a first pixel resolution, and the peripheral zone has a second set of image data and is rendered at a second pixel resolution lower than the first pixel resolution, and generating a control packet including rendering information associated with the foveal image. The method further includes time warping the foveal image to move at a viewer's position to form a time warped image, and sending the time warped image and the control packet to a video processor. The method further includes remapping the time warped image to a foveal region packed data block and a low resolution region packed data block, concatenating the control packet with the foveal region packed data block and the low resolution region packed data block to form a frame, and sending the frame to a display unit. The control packet is then parsed and decoded from the frame to obtain rendering information. The method also includes projecting a rendered display image according to the decoded rendering information.
[0010] According to some embodiments, a method of achieving a video pipeline implementation of binocular multiplexed foveated rendering with temporal warping includes rendering a first foveated image for a left eye including a first foveal zone and a first peripheral zone, where the first foveal zone has a first set of image data and is rendered at a first pixel resolution and the peripheral zone has a second set of image data and is rendered at a second pixel resolution lower than the first pixel resolution, and rendering a second foveated image for a right eye including a second foveal zone and a second peripheral zone, where the second foveal zone has a third set of image data and is rendered at a third pixel resolution and the peripheral zone has a fourth set of image data and is rendered at a fourth pixel resolution lower than the third pixel resolution. The method further includes generating a control packet including rendering information associated with the first foveated image and the second foveated image, time warping the first foveated image and the second foveated image to move to a viewer's position to form a first time warped image and a second time warped image, compressing the first time warped image and the second time warped image to form a first compressed image and a second compressed image, and sending the first compressed image, the second compressed image, and the control packet to a video processor.Next, the method further includes decompressing the first compressed image into a first reconstructed foveated image, decompressing the second compressed image into a second reconstructed foveated image, performing a second time warping on the first reconstructed foveated image and the second reconstructed foveated image based on the latest viewer pose data, remapping the first reconstructed foveated image into a first set of three separate color channels to form a first channel image, packing the first channel image with control packets to form a first frame, and transmitting the first frame to a first display for the left eye, wherein the first display for the left eye parses the control packets from the first frame and decodes the control packets to obtain rendering information for each of the three separate color channels for the first frame, and stores the rendering information of the first frame in a memory of the first display. For a right-eye display, the method includes remapping the second reconstructed foveated image into a second set of three separate color channels to form a second channel image, packing the second channel image into control packets to form a second frame, and transmitting the second frame to a second display for the right eye, where the second display for the right eye parses the control packets from the first frame and decodes the control packets to obtain rendering information for each of the three separate color channels for the second frame, and stores the rendering information for the second frame in a memory of the second display.
[0011] According to some embodiments, a method for implementing a video pipeline implementation of dynamically multiplexed foveated rendering with delay time warping and raster scan output includes rendering a foveated image including a foveal zone and a peripheral zone, where the foveal zone has a first set of image data and is rendered at a first pixel resolution, and the peripheral zone has a second set of image data and is rendered at a second pixel resolution lower than the first pixel resolution, and generating a control packet including rendering information associated with the foveal image. The method further includes time warping the foveal image to move at a viewer's position to form a time warped image, sending the time warped image and the control packet to a video processor, and performing delay time warping of the time warped image to form an updated image based on the latest viewer's posture data. The method also includes packing the updated image and the control packet to form a frame, and sending the frame to a display unit, where the display unit parses the control packet from the frame, decodes the control packet to obtain rendering information, and stores the rendering information of the frame in a memory of the display unit, and projecting the rendered display image according to the rendering information. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is an example image illustrating an implementation of foveated rendering (FR).
[0013] [Diagram 2] 2A-2E show some example artifacts that may be caused by subsampling in FR.
[0014] [Figure 3-1] 3A-3C are spatial profiles showing the visual field and foveal rendering using temporal multiplexing, according to some embodiments.
[0015] [Figure 3-2] 3D-3F are images illustrating native resolution, subsampling, and image blending according to an embodiment of the present invention.
[0016] [Figure 3-3] 3G-3H are text boxes illustrating sub-sampling and time multiplexing according to an embodiment of the present invention.
[0017] [Figure 4] FIG. 4 shows a simplified flowchart illustrating a method for generating foveated rendering using temporal multiplexing, according to some embodiments.
[0018] [Diagram 5] 5A-5D are spatial profiles showing left and right visual fields for foveated rendering using binocular multiplexing, according to some embodiments.
[0019] [Figure 6] 6A-6C are spatial profiles showing left and right visual fields for foveated rendering using binocular multiplexing in combination with temporal multiplexing, according to some embodiments.
[0020] [Figure 7] FIG. 7 illustrates a simplified flowchart illustrating a method for generating foveated rendering using binocular multiplexing, according to some embodiments.
[0021] [Figure 8] FIG. 8 shows a simplified flowchart illustrating a method for generating foveated rendering using a combination of temporal and binocular multiplexing, according to some embodiments.
[0022] [Figure 9]FIG. 9 illustrates an example control packet embedded in a video frame that provides FR information for use in a video pipeline according to some embodiments.
[0023] [Figure 10] FIG. 10 illustrates a block diagram of an example video pipeline for dynamically multiplexed FR, according to some embodiments.
[0024] [Figure 11] FIG. 11 illustrates a block diagram of an example video pipeline for dynamically multiplexed FR with temporal warping, according to some embodiments.
[0025] [Figure 12] FIG. 12 illustrates a block diagram of an example video pipeline for dynamically multiplexed foveated rendering including time warping configured for sequential color display, according to some embodiments.
[0026] [Figure 13] FIG. 13 illustrates a block diagram of an exemplary video pipeline for dynamically multiplexed foveated rendering with delay time warping and raster scan output, according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0027] Detailed Description of Specific Embodiments In foveated rendering (FR), a spatial profile of the field of view (FOV) is used to match rendering quality to the visual acuity of the human eye. The spatial profile includes a foveal zone and one or more peripheral regions. The foveal zone is rendered with maximum fidelity (e.g., at native pixel resolution), and the peripheral regions are rendered at a lower resolution. The spatial profile can be fixed within the FOV or can be shifted based on eye tracking data or content estimation. Due to inaccuracies and latency in eye tracking or content estimation, perceptual artifacts are often seen in FR.
[0028] According to some embodiments, a method of FR using temporal and / or binocular multiplexing of spatial profiles is provided. Such a method can improve the perceived visual quality of FR and reduce the visual artifacts of FR. The method can be used without eye tracking or can be used in combination with eye tracking. Such a method can be implemented in a video pipeline from a graphics processor to a display unit using a frame-specific control packet that provides information of FR for each frame. Thus, good visual quality of FR can be achieved while saving computation power and transmission bandwidth. These methods are described in more detail below.
[0029] FIG. 1 is an exemplary image showing an implementation of FR. A field of view (FOV) 110 is divided into two zones: a foveal zone 120 in a central region, and a peripheral zone 130 around the foveal zone 120. The foveal zone 120 is rendered at a native resolution (shown with a finer pixel grid), whereas the peripheral zone 130 is rendered at a reduced resolution (shown with a coarser pixel grid). The location of the foveal zone 120 within the FOV 110 can be determined by measuring fixations using an eye-tracking device, or by inferring where the viewer is looking based on the content. Alternatively, the location of the foveal zone 120 can be fixed, for example, to the center of the FOV 110, assuming that the viewer is looking at the center of the FOV 110.
[0030] The rendering pixel size of the peripheral zone 130 is often equal to an integer multiple of the pixel size of the foveal zone 120, for example through pixel binning or subsampling. For example, m×n native pixel groups can be replaced with one superpixel. In the example shown in FIG. 1, 2×2 native pixel groups are merged into one large pixel in the peripheral zone 130. Thus, the peripheral zone 130 has half the resolution in both the horizontal and vertical dimensions compared to the resolution of the foveal zone 120. This example will be used in the following description.
[0031] The assumption underlying FR is that the viewer sees the low-resolution region in his or her peripheral vision, where visual acuity is low enough that the effects of resolution reduction are imperceptible. In practice, this assumption can be invalidated due to eye-tracking errors or due to eye movements within a fixed FR. Thus, the viewer may see some artifacts due to subsampling in FR.
[0032] 2A-2E show some example artifacts that may be caused by subsampling. Assuming 2×2 subsampling, a 1-pixel-wide line 210 shown in FIG. 2A becomes a 2-pixel-wide line 220 shown in FIG. 2B. That is, the maximum spatial frequency is halved. Because adjacent pixels are composited, their luminance is averaged. Thus, contrast is reduced, as shown in FIGS. 2A and 2B. Subsampling may also change color. For example, the red and green lines 230 and 240 at the border between the red and green areas may be merged into a wide yellow line 250, as shown in FIGS. 2C and 2D. Subsampling may also reduce the readability of fine text, as shown in FIG. 2E, which shows text in native pixels on the left and text in subsampled pixels on the right.
[0033] The above artifacts can be more noticeable when the content moves as a result of the viewer's head movement or the content itself. For example, a boundary can be seen between the foveal zone and the peripheral zone, as indicated by differences in luminance and / or contrast. The above artifacts can be more noticeable when the content has high contrast, for example, white text on a black background viewed by a VR headset, or white text against a dark background viewed by an AR headset.
[0034] According to some embodiments, methods are provided for improving the perceptual quality of images rendered by FR by reducing the conspicuousness of the above-mentioned artifacts. Instead of using a fixed spatial profile for rendering (i.e., foveal native resolution vs. peripheral low resolution), a temporal and / or binocular multiplexing of different spatial profiles is used for rendering. These methods are also referred to herein as temporal "dithering". These solutions can have the advantage of being computationally lightweight while not requiring accurate and fast eye tracking.
[0035] 3A-3C are spatial profiles illustrating a field of view and foveal rendering using temporal multiplexing, according to some embodiments. FIG. 3A shows a first spatial profile in which the FOV 310 is divided into a first foveal zone 320 (represented by dark pixels) and a first peripheral zone 330 (represented by white pixels). FIG. 3B shows a second spatial profile in which the FOV 310 is divided into a second foveal zone 340 and a second peripheral zone 350. As shown, the second foveal zone 340 is shifted horizontally (e.g., by 4 native pixels in the X direction) relative to the first foveal zone 320.
[0036] According to some embodiments, the first spatial profile and the second spatial profile are temporally multiplexed in a series of frames. For example, the first spatial profile can be used to render odd frames, and the second spatial profile can be used to render even frames. In this way, the foveal zone is dynamically moved from frame to frame in a series of frames. Thus, for regions of the FOV 310 where the first foveal zone 320 and the second foveal zone 340 overlap (e.g., the dark center column shown in FIG. 3C), an image rendered at native resolution is always presented. For regions of the first foveal zone 320 that do not overlap with the second foveal zone 340, or for regions of the second foveal zone 340 that do not overlap with the first foveal zone 320 (e.g., the gray column shown in FIG. 3C), native resolution images and subsampled images are presented alternately from frame to frame.
[0037] Assuming the display has a sufficiently high refresh rate (e.g., 120Hz), the native resolution image and the subsampled image can be blended together as perceived by the viewer. Blending the native resolution image and the subsampled image can help restore high spatial frequencies and luminance contrast in the viewer's visual perception. Figures 3D-3F show examples.
[0038] 3D-3F are images illustrating native resolution, subsampling, and image blending according to an embodiment of the present invention. FIG. 3D shows a native resolution image including a one-pixel-wide line 360. FIG. 3E shows a subsampled image in which the one-pixel-wide line 360 becomes a two-pixel-wide line 370. FIG. 3F shows the result of blending the two images shown in FIG. 3D and FIG. 3E as may be perceived by a viewer. As shown in FIG. 3F, the high spatial resolution and contrast of the native resolution image is somewhat restored.
[0039] 3G-3H are text boxes illustrating subsampling and temporal multiplexing according to an embodiment of the present invention. FIG. 3G shows some text in a subsampled image. FIG. 3H shows text multiplexed with a native resolution image and a subsampled image. As shown, the readability of the text in FIG. 3H is improved compared to FIG. 3G. Thus, in the example shown in FIG. 3A-3C, by temporally multiplexing the first spatial profile shown in FIG. 3A and the second spatial profile shown in FIG. 3B, the effective foveal zone (e.g., the combination of the dark and gray regions in FIG. 3C) can be enlarged compared to the foveal zone 320 or 340 of each individual spatial profile.
[0040] According to various embodiments, the location of the foveal zone can be spatially shifted between successive frames in the horizontal direction (e.g., X direction), or vertical direction (e.g., Y direction), or both directions (e.g., a combination of X and Y directions). Furthermore, the direction and amount of spatial shifting can be dynamically changed. The frame rate can be limited by the capabilities of the display (e.g., spatial light modulator or SLM). For example, the frame rate can be 120 Hz or higher. In some embodiments, the foveal zone can be spatially shifted to a set of predefined positions in a fixed or random order to cover as much of the FOV 310 as possible. Thus, the viewer may perceive a high quality image across the entire FOV, even if the viewer's line of sight changes.
[0041] FIG. 4 shows a simplified flowchart illustrating a method 400 for generating foveated rendering using temporal multiplexing, according to some embodiments.
[0042] The method 400 includes generating a first spatial profile of the FOV by dividing the FOV into a first foveal zone and a first peripheral zone, at 402. The first foveal zone is rendered at a first pixel resolution and the first peripheral zone is rendered at a second pixel resolution that is lower than the first pixel resolution.
[0043] The method 400 further includes generating a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone, at 404. The second foveal zone is spatially offset from the first foveal zone. The second foveal zone is rendered at a first pixel resolution and the second peripheral zone is rendered at a second pixel resolution.
[0044] The method 400 further includes, at 406, temporally multiplexing the first spatial profile and the second spatial profile in a series of frames such that a viewer, when rendered at the first pixel resolution, perceives an image rendered in an area of the first foveal zone that does not overlap with the second foveal zone and / or an area of the second foveal zone that does not overlap with the first foveal zone.
[0045] It should be appreciated that the particular steps illustrated in FIG. 4 provide a particular method of generating a foveated rendering according to some embodiments. Other sequences of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, each step illustrated in FIG. 4 may include multiple sub-steps that may be performed in various orders as appropriate for the individual step. Additionally, additional steps may be added and some steps may be removed depending on the particular application. Those skilled in the art will recognize many variations, modifications, and alternatives.
[0046] According to some embodiments, binocular multiplexing can be applied additionally or alternatively to reduce perceptual artifacts. Figures 5A-5D show spatial profiles illustrating left and right visual fields for foveated rendering using binocular multiplexing according to some embodiments. Referring to Figure 5A, in a first spatial profile, the field of view (FOV) of the left eye 510 is divided into a first foveal zone 520 (represented by gray pixels) and a first peripheral zone 530 (represented by white pixels). In a second spatial profile, the FOV of the right eye 540 is divided into a second foveal zone 550 and a second peripheral zone 560. It is assumed that the FOV for the left eye 510 and the FOV for the right eye 540 are identical. As shown, the first foveal zone 520, instead of being fixed at the center of the FOV, is shifted to the right and the second foveal zone 550 is shifted to the left.
[0047] With reference to Figure 5B, if the viewer's fixation reaches the center of the FOV where the first foveal zone 520 and the second foveal zone 550 overlap, the viewer may see a natural resolution image to both eyes in his / her central vision. With reference to Figure 5A, if the viewer's fixation reaches the right side of the FOV, the viewer may see a native resolution image to the left eye and a subsampled image to the right eye in his / her central vision. With reference to Figure 5C, if the viewer's fixation is located on the left side of the FOV, the viewer may see a native resolution image to the right eye and a subsampled image to the left eye in his / her central vision.
[0048] FIG. 5D shows the effective spatial profile. When viewing the area represented by dark pixels, the viewer may see a native resolution image in both eyes. When viewing the area represented by gray pixels, the viewer may see a native resolution image in only one eye. As shown in FIG. 5D, the combined area (represented by gray and dark pixels) where the native resolution image is seen by at least one eye is larger than the foveal zone 520 or 550 in each individual spatial profile. Therefore, subsampling artifacts can be suppressed. According to various embodiments, the position of the foveal zone 520 and 550 can be shifted horizontally (e.g., in the X direction), or vertically (e.g., in the Y direction), or both directions (e.g., a combination of the X and Y directions).
[0049] According to some embodiments, binocular multiplexing can be combined with temporal multiplexing. FIGS. 6A-6C are spatial profiles showing left and right fields of view for foveated rendering using binocular multiplexing in combination with temporal multiplexing according to some embodiments. With reference to FIG. 6A, in a first frame, a first spatial profile of the left FOV can have a foveal zone 610 shifted to the right, and a second spatial profile of the right FOV can have a foveal zone 620 shifted to the left. With reference to FIG. 6B, in a second frame, a third spatial profile of the left FOV can have a foveal zone 630 shifted to the left, and a fourth spatial profile of the right FOV can have a foveal zone 640 shifted to the right. In some embodiments, in a series of frames, the first frame as shown in FIG. 6A can be every odd frame, and the second frame as shown in FIG. 6B can be every even frame.
[0050] Referring to FIG. 6C, in the third frame, the fifth spatial profile of the left FOV can have a foveal zone 650 shifted upward, and the sixth spatial profile of the right FOV can have a foveal zone 660 shifted downward. In some embodiments, the direction and amount of spatial shift can be dynamically changed. For example, the foveal zones of both the left and right FOVs can be spatially shifted to a set of predefined positions in a fixed or random order to cover as much of the FOV as possible. Thus, the viewer may perceive a high quality image across the entire FOV while saving significant bandwidth. The movement of the foveal zone can be calculated using the minimum and maximum possible interpupillary distance (IPD) to ensure that good visual results can be achieved for the target viewer.
[0051] According to some embodiments, the method of temporal and binocular multiplexing of spatial profiles can be applied to various types of FR implementations, including, for example, fixed FR, FR with eye tracking, or content-based FR. When applied to FR with eye tracking, the method described herein can effectively extend the foveal region and thus reduce artifacts generated by inaccurate eye tracking. When applied to content-based FR, the method described herein can reduce artifacts caused by prediction errors. When applied to fixed FR, the method described herein can help smooth the transition of visual quality from the highest resolution in the foveal region to multiple resolutions near the periphery and sub-sample resolutions in the periphery.
[0052] FIG. 7 shows a simplified flowchart illustrating a method 700 for generating foveated rendering using binocular multiplexing, according to some embodiments.
[0053] The method 700 includes generating 702 a first spatial profile of the FOV by dividing the FOV into a first foveal zone and a first peripheral zone, the first foveal zone being rendered at a first pixel resolution and the first peripheral zone being rendered at a second pixel resolution that is lower than the first pixel resolution.
[0054] The method 700 further includes generating 704 a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone, the second foveal zone being spatially offset from the first foveal zone, the second foveal zone being rendered at a first pixel resolution and the second peripheral zone being rendered at a second pixel resolution.
[0055] The method 700 further includes, at 706, multiplexing the first spatial profile and the second spatial profile for each of the left and right eyes of the viewer such that, when rendered at the first pixel resolution, the viewer perceives an image rendered in an area of the first foveal zone that does not overlap with the second foveal zone and / or an area of the second foveal zone that does not overlap with the first foveal zone.
[0056] It should be appreciated that the particular steps illustrated in FIG. 7 provide a particular method of generating a foveated rendering according to some embodiments. Other sequences of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, individual steps illustrated in FIG. 7 may include multiple sub-steps that may be performed in various orders as appropriate for the individual step. Additionally, additional steps may be added and some steps may be removed depending on the particular application. Those skilled in the art will recognize numerous variations, modifications, and alternatives.
[0057] FIG. 8 shows a simplified flowchart illustrating a method 800 for generating foveated rendering using a combination of temporal and binocular multiplexing, according to some embodiments.
[0058] The method 800 includes generating a first spatial profile of the FOV by dividing the FOV into a first foveal zone and a first peripheral zone, at 802. The first foveal zone is rendered at a first pixel resolution and the first peripheral zone is rendered at a second pixel resolution that is lower than the first pixel resolution.
[0059] The method 800 further includes generating a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone, at 804. The second foveal zone is spatially offset from the first foveal zone. The second foveal zone is rendered at a first pixel resolution and the second peripheral zone is rendered at a second pixel resolution.
[0060] The method 800 further includes generating a third spatial profile of the FOV by dividing the FOV into a third foveal zone and a third peripheral zone, at 806. The third foveal zone is spatially offset from the first foveal zone. The third foveal zone is rendered at the first pixel resolution and the third peripheral zone is rendered at the second pixel resolution.
[0061] The method 800 further includes generating a fourth spatial profile of the FOV by dividing the FOV into a fourth foveal zone and a fourth peripheral zone, at 808. The fourth foveal zone is spatially offset from the third foveal zone. The fourth foveal zone is rendered at the first pixel resolution and the fourth peripheral zone is rendered at the second pixel resolution.
[0062] The method 800 further includes, at 810, multiplexing the first and second spatial profiles for the left and right eyes of the viewer, respectively, at odd frames.
[0063] The method 800 further includes, at 812, multiplexing the third and fourth spatial profiles for the left and right eyes of the viewer, respectively, in even frames.
[0064] It should be appreciated that the particular steps illustrated in FIG. 8 provide a particular method of generating a foveated rendering according to some embodiments. Other sequences of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, individual steps illustrated in FIG. 8 may include multiple sub-steps that may be performed in various orders as appropriate for the individual step. Additionally, additional steps may be added and some steps may be removed depending on the particular application. Those skilled in the art will recognize many variations, modifications, and alternatives.
[0065] The FR rendering methods described herein can be implemented in hardware and / or software pipelines in a variety of ways. For example, they can be implemented in the design of an AR / VR system via firmware or middleware. Such implementations can be transparent to the application. Alternatively, they can be implemented in existing AR / VR systems via operating system (OS) software or individual applications.
[0066] According to some embodiments, resource savings can be realized throughout the rendering to display pipeline. A graphics driver (e.g., a graphics processing unit or GPU) can first generate both low-resolution and high-resolution images for each frame and pack them to minimize the video payload. The positions of the low-resolution and high-resolution images can be assumed to change from frame to frame. A control packet can be embedded within the video frame that provides frame-specific information of the FR, including pixel indexing of the foveal zone. The information provided by the control packet can assist the display unit (e.g., an SLM ASIC) in unpacking the image data for each frame.
[0067] FIG. 9 illustrates an exemplary control packet according to some embodiments. The control packet can be embedded in the video frame generated by the GPU. The control packet can include information such as whether FR mode is enabled, the downsampling ratio (e.g., 4:1, 9:1, 16:1, etc.), the starting row and starting column index of the foveal region, etc. The control packet can serve as a map for the display unit (e.g., by the SLM ASIC) to unpack the image data in the video frame. The downsampling ratio can be dynamically changed on a frame-by-frame basis. For example, the ratio can be 16:1 for most frames and changed to 4:1 for those frames with text content.
[0068] In some embodiments, each video frame can include three primary (e.g., red, green, and blue) color channels. Each color channel can be enabled or disabled in FR mode independently. For example, since the human eye has the strongest perception of green, it may be advantageous to have full resolution green content throughout the frame and apply FR only to the red and blue content. The foveal zone of each color channel can have its own ratio of downsampling, size, and location with starting row and starting column index. For example, the green content can have a lower ratio of downsampling than the red and blue colors. Also, the foveal zones of the three color channels do not need to be aligned with each other.
[0069] It should be understood that the information that may be included in a control packet is not limited to the specific information described above. Any information or factor that may affect the rendering, processing, display, etc. of an image may be included in a control packet.
[0070] 10 illustrates a block diagram of an example video pipeline implementation of dynamic multiplexed foveated rendering, according to some embodiments. An image 1020 (e.g., a video frame) may be rendered in a GPU 1010. The image 1020 includes a foveal zone 1022 rendered at a high resolution and a peripheral zone 1024 rendered at a low resolution. The image data (including the high resolution image data of the foveal zone 1022 and the low resolution image data of the peripheral zone 1024) is packed into a video frame 1030. For example, the high resolution image data may be packed into a first image block 1034, and the low resolution image data may be packed into a second image block 1036.
[0071] The control packet 1032 is then concatenated with the first image block 1034 and the second image block 1036. The control packet 1032 may include information regarding FR rendering (e.g., as shown in FIG. 9). The video frame 1030 may be transmitted to a display unit 1050 (e.g., an SLM ASIC) via a channel link 1040. Packing the video frame with the control packet 1032 may significantly reduce the required payload and data rate, thereby minimizing the bandwidth and power on the channel link 1040.
[0072] The display unit 1050 can parse the control packets 1032 from the first image block 1034 and the second image block 1036 using a frame analyzer 1060 (e.g., a decoder). For example, the control packet 1032 can be in the first row (e.g., row 0) of the video frame 1030. The control packet decoder 1070 can decode the control packet 1032. The information provided by the control packet 1032 can then be used to map the image data in the first image block 1034 and the second image block 1036 to the foveal zone and the peripheral zone in the video memory 1080, respectively. For low-resolution areas, the large pixels can be mapped to a few native pixels of the display (e.g., to four native pixels if the subsampling ratio is 4:1). The display unit 1050 performs this decoding process frame by frame. The display unit 1050 then projects the images stored in the video memory 1080 to a viewer (e.g., outputting photons via the SLM). In some embodiments, timestamps may also be included in the video frames 1030 to aid in synchronization and latency management, as well as partial screen refresh tasks, blank modes, and the like.
[0073] According to some embodiments, the GPU renders the foveated image and generates the associated control packets. A time warping function can be performed on the image data based on the latest pose prediction from the computer vision processor to provide an image to the display unit to reflect the viewer's viewpoint based on the latest pose prediction. According to some embodiments, the time warping can be performed before sending the image data and control packets to the display unit. As shown in FIG. 11, the time warping can be performed in the time warping block 1140 before the image data and control packets are sent to the video processor 1150, as well as in the time warping remapping block 1152 of the video processor 1150. Thus, in some embodiments, the time warping can be performed by a computer vision processor coupled to the display unit. The foveated image is then sent to the wearable video processor, and the control packets are sent over a secondary data channel. The video processor then performs a delay time warping (e.g., using the latest pose data from the sensor suite of the wearable device) and reformats the foveated image with the foveal and low-resolution regions. This method can work well for global refresh SLM, which can unpack the image before display. For a color sequential display, red, green and blue packed images can be generated and transmitted separately.
[0074] 11 shows a block diagram of an example video pipeline for dynamically multiplexed foveated rendering with time warping, according to some embodiments. A GPU 1110 generates a foveated image 1120 and associated control packets 1130 (e.g., as shown in FIG. 9). A time warping block 1140 performs a time warping function on the foveated image 1120 to account for motion at the viewer's position. The image data and control packets 1130 are then sent to the wearable device's video processor 1150 via a headset link.
[0075] In the video processor 1150, the time warping and remapping block 1152 can perform delay time warping using the latest pose data (e.g., from the sensor suite of the wearable device). For example, if there is a large head movement, the latest pose data can be used to update the boundary. The time warping and remapping block 1152 can also remap the foveal image 1120 to a foveal region data block 1154 and a low resolution region data block 1156. The control packet 1130 can be concatenated with the foveal region data block 1154 and the low resolution region data block 1156 to form a video frame and sent to the display unit 1160 (e.g., an SLM ASIC). The display unit 1160 can use the control packet to unpack the video frame, similar to the display unit 1050 shown in FIG. 10 described above.
[0076] FIG. 12 illustrates a block diagram of an exemplary video pipeline for dynamic multiplexed foveated rendering including time warping configured for sequential color display, according to some embodiments. The GPU 1210 renders a first foveated image 1212 for the left eye and a second foveated image 1214 for the right eye. An associated control packet 1216 is generated that provides FR information (e.g., as shown in FIG. 9) for both the first foveated image 1212 and the second foveated image 1214. The time warping block 1218 performs a time warping function on the first foveated image 1212 and the second foveated image 1214. The compression block 1219 performs compression of the image data. The compressed image data and the control packet 1216 are then transmitted to the video processor 1220 of the wearable device via a headset link.
[0077] In the video processor 1220, a decompression block 1222 decompresses the image data and recovers a first foveated image 1212 and a second foveated image 1214. A first time warping and remapping block 1224 performs time warping on the first foveated image 1212 based on the latest pose data. The first time warping and remapping block 1224 also maps the first foveated image 1212 into three separate color channels (e.g., red 1232a / 1232b, green 1234a / 1234b, and blue 1236a / 1236b). The three color channels (i.e., 1232a, 1234a, and 1236a) are packed together with a control packet as a first video frame 1228 to be transmitted to a first display 1230 (e.g., an LCOS display) for the left eye. A second time warping and remapping block 1226 performs time warping on the second foveated image 1214 based on the latest pose data. The second time warping and remapping block 1226 also maps the second foveated image 1214 into three separate color channels (i.e., 1232b, 1234b, and 1236b). The three color channels, along with a control packet, are packed as a second video frame 1229 that is sent to a second display 1240 (e.g., an LCOS display) for the right eye.
[0078] The first display 1230 can unpack the first video frame 1228 using the control packets in a manner similar to the display unit 1050 shown in FIG. 10 described above. In this case, image data for each of the three color channels 1232a, 1234a, and 1236a is stored in a video memory and projected to the viewer's left eye. As described above, each of the three color channels 1232a, 1234a, and 1236a can have its independent foveal zone, downsampling ratio, etc. The foveal zones of the three color channels 1232a, 1234a, and 1236a do not have to be aligned with each other. In some embodiments, the three color channels 1232a, 1234a, and 1236a can be displayed sequentially to the viewer. The second display 1240 can unpack the second video frame 1229 in a similar manner.
[0079] For rolling shutter type displays, image data may be packed differently to keep it in a rasterized format. FIG. 13 shows a block diagram of an example video pipeline for dynamic multiplexed foveated rendering with delayed time warping and raster scan output, according to some embodiments. A foveated image 1312 is rendered in a GPU 1310. A time warping block 1314 performs a time warping function on the foveated image 1312. For the foveated image 1312, a control packet 1316 can be created that includes information about FR rendering (e.g., as shown in FIG. 9). The image data and control packet 1316 are transmitted to a video processor 1320 of the wearable device via a headset link.
[0080] In the video processor 1320, a time warping and remapping block 1324 can perform delay time warping on the foveated image using the latest pose data (e.g., from the wearable device's sensor suite). The time warped image and control packets 1316 are packaged as a video frame 1322 and then sent to a display unit 1330.
[0081] In the display unit 1330, the frame analyzer 1334 can parse the control packets 1316 from the image data 1312. The control packet decoder 1336 can decode the control packets 1316 that accompany the image data 1312. The information provided by the control packets 1316 can then be used to map the image data 1312 to the video memory 1332. For example, a large pixel in a low-resolution region can be mapped to a few native pixels of the display (e.g., to four native pixels if the subsampling ratio is 4:1). The display unit 1050 can then project the image stored in the video memory 1332 to a viewer (e.g., outputting photons via an SLM). Because foveated rendering can be dynamically changed on a frame-by-frame basis, the display unit 1330 uses the FR information provided in the control packets 1316 on a frame-by-frame basis to ensure correct mapping.
[0082] For rolling shutter type displays, image data 1312 is maintained in rasterized form throughout the pipeline so that image data can be scanned and output as it is scanned. There is no need to wait for an entire frame to be received. Thus, video memory 1332 can be a relatively small line buffer to send out newly arrived image data. A large buffer to hold an entire frame is not needed. Also, latency can be kept relatively low.
[0083] It is also understood that the examples and embodiments described herein are for illustrative purposes only, and that various changes or modifications in light thereof will be suggested to those skilled in the art and are to be included within the spirit and scope of this application and the scope of the appended claims.
Claims
1. A method for generating a foveated rendering for a display, the method comprising: generating a first spatial profile of a field of view (FOV) of the display by dividing the FOV into a first foveal zone and a first peripheral zone, the first foveal zone including first pixels rendered at a first pixel resolution, the first peripheral zone including second pixels rendered at a second pixel resolution lower than the first pixel resolution, the first foveal zone including a first region including a subset of the first pixels and a second region including the remainder of the first pixels; generating a second spatial profile of the FOV of the display by dividing the FOV into a second foveal zone and a second peripheral zone, the second foveal zone being spatially offset from the first foveal zone, the second foveal zone including third pixels rendered at the first pixel resolution, the second peripheral zone including fourth pixels rendered at the second pixel resolution, the second foveal zone including a third region including a subset of the third pixels and a fourth region including the remainder of the third pixels; temporally multiplexing the first spatial profile and the second spatial profile over a series of frames such that a viewer perceives an image rendered in the first region of the first foveal zone that does not overlap with the second foveal zone and / or the third region of the second foveal zone that does not overlap with the first foveal zone; A method comprising:
2. The method of claim 1 , wherein the first foveal zone and the second foveal zone partially overlap each other.
3. The method of claim 1 , wherein the series of frames has a frame frequency of approximately 120 Hz.
4. The method of claim 1 , wherein the ratio between the first pixel resolution and the second pixel resolution is 2:1 in each of two orthogonal directions.
5. 2. The method of claim 1, wherein the second foveal zone is spatially offset from the first foveal zone in a first direction, a second direction orthogonal to the first direction, or in both the first direction and the second direction.
6. The method of claim 5 , wherein the spatial offset between the second foveal zone and the first foveal zone is dynamically changed in a series of frames.
7. The method of claim 6 , wherein the dynamic change of the spatial offset between the second foveal zone and the first foveal zone follows a pattern.
8. 2. The method of claim 1 , wherein each of the first spatial profile and the second spatial profile includes three sub-spatial profiles for each of three primary colors, and at least one of the three sub-spatial profiles is rendered at the second pixel resolution for the entire FOV.
9. The method of claim 1 , wherein the first foveal zone or the second foveal zone is set at a predetermined position within the FOV of the display.
10. The method of claim 9 , wherein the predetermined location of the first foveal zone or the second foveal zone is determined based on measurements of a viewer's eye position and eye movement.
11. A method for generating a foveated rendering for a display, comprising: generating a first spatial profile of a field of view (FOV) of the display by dividing the FOV into a first foveal zone and a first peripheral zone, the first foveal zone including first pixels rendered at a first pixel resolution, the first peripheral zone including second pixels rendered at a second pixel resolution lower than the first pixel resolution, the first foveal zone including a first region including a subset of the first pixels and a second region including the remainder of the first pixels; generating a second spatial profile of the FOV of the display by dividing the FOV into a second foveal zone and a second peripheral zone, the second foveal zone being spatially offset from the first foveal zone, the second foveal zone including third pixels rendered at the first pixel resolution, the second peripheral zone including fourth pixels rendered at the second pixel resolution, the second foveal zone including a third region including a subset of the third pixels and a fourth region including the remainder of the third pixels; multiplexing the first spatial profile and the second spatial profile for each of the left and right eyes of the viewer viewing the display such that the viewer perceives an image rendered in the first region of the first foveal zone that does not overlap with the second foveal zone and / or the third region of the second foveal zone that does not overlap with the first foveal zone; A method comprising:
12. The method of claim 11 , wherein the first foveal zone and the second foveal zone partially overlap each other.
13. The method of claim 11 , wherein the series of frames has a frame frequency of about 120 Hz.
14. The method of claim 11 , wherein the ratio between the first pixel resolution and the second pixel resolution is 2:1 in each of two orthogonal directions.
15. 12. The method of claim 11, wherein the second foveal zone is spatially offset from the first foveal zone in a first direction, a second direction orthogonal to the first direction, or in both the first direction and the second direction.
16. 12. The method of claim 11 , wherein each of the first spatial profile and the second spatial profile includes three sub-spatial profiles for each of three primary colors, and at least one of the three sub-spatial profiles is rendered at the second pixel resolution for the entire FOV.
17. The method of claim 11 , wherein the first foveal zone or the second foveal zone is set at a predetermined position within the FOV of the display.
18. The method of claim 11 , wherein the spatial offset between the second foveal zone and the first foveal zone is dynamically changed over a series of frames.
19. The method of claim 18 , wherein the dynamic variation of the spatial offset between the second foveal zone and the first foveal zone follows a pattern.
20. 1. A method for generating a foveated rendering, the method comprising: generating a first spatial profile of a field of view (FOV) by dividing the FOV into a first foveal zone and a first peripheral zone, the first foveal zone being rendered at a first pixel resolution and the first peripheral zone being rendered at a second pixel resolution lower than the first pixel resolution; generating a second spatial profile of the FOV by dividing the FOV into a second foveal zone and a second peripheral zone, the second foveal zone being spatially offset from the first foveal zone, the second foveal zone being rendered at the first pixel resolution and the second peripheral zone being rendered at the second pixel resolution; generating a third spatial profile of the FOV by dividing the FOV into a third foveal zone and a third peripheral zone, the third foveal zone being spatially offset from the first foveal zone, the third foveal zone being rendered at the first pixel resolution and the third peripheral zone being rendered at the second pixel resolution; generating a fourth spatial profile of the FOV by dividing the FOV into a fourth foveal zone and a fourth peripheral zone, the fourth foveal zone being spatially offset from the third foveal zone, the fourth foveal zone being rendered at the first pixel resolution and the fourth peripheral zone being rendered at the second pixel resolution; multiplexing the first spatial profile and the second spatial profile for the left eye and the right eye of a viewer, respectively, in odd frames; multiplexing the third spatial profile and the fourth spatial profile for the left eye and the right eye of a viewer, respectively, in even frames; A method comprising:
21. 21. The method of claim 20, wherein the third foveal zone is spatially offset from the first foveal zone in a first direction, and the fourth foveal zone is spatially offset from the second foveal zone in a second direction opposite the first direction.
22. 22. The method of claim 21, wherein the spatial offset between the first foveal zone and the third foveal zone is dynamically changed in a series of frames.
23. 22. The method of claim 21, wherein the spatial offset between the second foveal zone and the fourth foveal zone is dynamically changed in a series of frames.
24. 21. The method of claim 20, wherein each of the first spatial profile, the second spatial profile, the third spatial profile, and the fourth spatial profile includes three sub-spatial profiles for each of three primary colors, and at least one of the three sub-spatial profiles is rendered at the second pixel resolution for the entire FOV.
25. 23. The method of claim 22, wherein the dynamic variation of the spatial offset between the first foveal zone and the third foveal zone follows a pattern.
26. 24. The method of claim 23, wherein the dynamic variation of the spatial offset between the second foveal zone and the fourth foveal zone follows a pattern.