Gaze-based super-resolution for augmented reality devices
By determining the user's gaze characteristics in extended reality devices and using machine learning models for super-resolution reconstruction, the image resolution problem caused by low-resolution cameras is solved, improving user experience and reducing system power consumption.
Patent Information
- Application Number
- CN202480005775.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-01-20
- Publication Date
- 2025-07-29
AI Technical Summary
Low-resolution cameras lead to adverse effects on image content resolution in extended real-life devices, affecting user experience.
The user's gaze characteristics are determined through the computing device, a fovea map is generated, and a machine learning model is used for super-resolution reconstruction to improve image resolution.
Improves the resolution of images in extended real-life devices, enhances the immersive experience for users, and reduces system power consumption and cost.
Smart Images

Figure CN120390918A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to extended reality devices, and more particularly, to foveation-based super-resolution techniques for extended reality devices. Background Art
[0002] Extended reality (XR) environments can generally include: a real-world environment including XR content that covers one or more features of the real-world environment; and / or a fully immersive virtual environment in which a user can navigate and experience using one or more avatars. In a typical XR system, image data can be rendered on, for example, a robust head-mounted display (HMD) that can be coupled to a basic graphics generation device responsible for generating the image data via a physical wired connection or a wireless connection. To reduce power consumption and system cost, in some instances, the XR system can include a low-resolution camera for capturing the real-world environment. However, the low-resolution camera may have an adverse effect on the resolution of the image content. Therefore, it may be useful to provide techniques for improving XR systems. Summary of the Invention
[0003] Current embodiments include techniques for providing foveation-based super-resolution for image content rendered by an extended reality (XR) device.
[0004] According to a first aspect, a method is provided that includes: by a computing device configured to be worn by a user: rendering an extended reality (XR) environment on one or more displays of the computing device; determining a context of the XR environment relative to the user, where determining the context includes determining one or more characteristics associated with at least one eye of the user relative to the content displayed within the XR environment; generating a fovea map based on the one or more characteristics associated with at least one eye of the user, where the fovea map includes a plurality of fovea regions, and where the plurality of fovea regions includes a plurality of partitions, each partition corresponding to a low-resolution region of the content of the corresponding partition; inputting one or more of the plurality of partitions into a machine learning model that is trained to generate a super-resolution reconstruction of the fovea map based on regions of interest identified within one or more of the plurality of partitions; and outputting, by the machine learning model, the super-resolution reconstruction of the fovea map.
[0005] The method may further include: causing the XR display device to display one or more image frames using the super-resolution reconstruction of the fovea map.
[0006] The method may further include: determining the context of the XR environment relative to the user based on a prediction of the position of at least one eye of the user relative to the content over a plurality of image frames to be displayed next.
[0007] Determining the context of the XR environment relative to the user may include determining one or more of the following: the head pose of the user; the eye gaze of the user; or the depth of eye vector convergence.
[0008] The machine learning model may also be trained to: generate a multi-frame temporal super-resolution reconstruction of the fovea map based on the region of interest identified within one or more of the plurality of partitions. The multi-frame temporal super-resolution reconstruction may be applied to reconstruct the texture details of the fovea map.
[0009] One or more of the plurality of partitions may correspond to a central fovea region of the user's approximately 30°×30° field of view (FOV).
[0010] The machine learning model may include a convolutional neural network (CNN), a deep convolutional neural network (DCNN), a vision transformer (ViT), or a generative adversarial network (GAN).
[0011] According to a second aspect, there is provided a computing device configured to be worn by a user, the computing device including: one or more displays; one or more non-transitory computer-readable storage media including a plurality of instructions; and one or more processors coupled to the one or more displays, the one or more processors configured to execute the above instructions to perform the method according to the first aspect.
[0012] According to a third aspect, there is provided a computer-readable medium including a plurality of instructions which, when executed by one or more processors of a computing device configured to be worn by a user, cause the computing device to perform the method according to the first aspect. The medium may be non-transitory. According to a fourth aspect, there is provided a computer program product including a plurality of instructions which, when executed by one or more processors of a computing device configured to be worn by a user, cause the computing device to perform the method according to the first aspect.
[0013] Certain other embodiments may include all, some, of the multiple components, elements, features, functions, operations, or steps of the embodiments disclosed above, or may not include the multiple components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments in accordance with the present invention are particularly disclosed in the appended claims directed to methods, storage media, systems, and computer program products, wherein any feature mentioned in one claim category (e.g., method) may also be claimed in another claim category (e.g., system). The dependencies or back-references in the appended claims are chosen only for formality reasons. However, any subject matter resulting from intentionally back-referencing any previous claim (especially multiple dependent claims) may also be claimed, such that any combination of the multiple claims and their multiple features is disclosed and may be claimed, regardless of the dependencies chosen in the appended claims. The subject matter that may be claimed includes not only combinations of the multiple features stated in the appended claims, but also any other combinations of the multiple features in the claims, wherein each feature mentioned in the claims may be combined with any other feature or combination of features in the claims. Additionally, any of the multiple embodiments and features described or depicted herein may be claimed in a separate claim, and / or in any combination with any embodiment or feature described or depicted herein or in any combination with any feature in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1A An example of an extended reality (XR) display device is shown.
[0015] Figure 1B An example of a passthrough feature is shown.
[0016] Figure 2A and Figure 2B An example sensor readout of an image sensor is shown.
[0017] Figure 3 An example image divided into multiple different partitions is shown.
[0018] Figure 4A and Figure 4B An example process of foveated region processing is shown.
[0019] Figure 5 A gaze-based super-resolution reconstruction model is shown, which can be adapted to provide gaze-based super-resolution for image content rendered by an extended reality (XR) device.
[0020] Figure 6A and Figure 6BOne or more image reconstruction examples 600A and 600B are shown respectively.
[0021] Figure 7 A flowchart of a method for providing gaze-based super resolution for image content rendered for an extended reality (XR) device is shown.
[0022] Figure 8 An example computer system is shown. Detailed Description
[0023] An extended reality (XR) environment can generally include: a real-world environment including XR content that covers one or more features of the real-world environment; and / or a fully immersive virtual environment in which a user can navigate and experience using one or more user avatars. In a typical XR system, image data can be rendered on, for example, a robust head-mounted display (HMD) that can be coupled to a basic graphics generation device responsible for generating the image data via a physical wired connection or a wireless connection. To reduce power consumption and system cost, in some instances, an XR system can include a low-resolution camera for capturing the real-world environment. However, the low-resolution camera may have an adverse effect on the resolution of the image content. Therefore, it may be useful to provide techniques for improving XR systems.
[0024] Accordingly, the present embodiment includes techniques for providing gaze-based super resolution for image content rendered by an extended reality (XR) device. In certain embodiments, a computing device configured to be worn by a user can render an extended reality (XR) environment on one or more displays of the computing device. In certain embodiments, the computing device can then determine the context of the XR environment relative to the user. For example, in one embodiment, the computing device can determine the context by determining one or more characteristics associated with at least one eye of the user relative to the content displayed within the XR environment. In certain embodiments, the computing device can determine the context of the XR environment relative to the user by determining one or more of the following: the head pose of the user; or the eye gaze of the user. In certain embodiments, the computing device can determine the context of the XR environment relative to the user based on a prediction of the position of at least one eye of the user relative to the content over the next few image frames to be displayed. In another embodiment, the computing device can determine the context of the XR environment relative to the user based on the depth of eye vector convergence.
[0025] In some embodiments, the computing device may then generate a fovea map based on one or more characteristics associated with at least one eye of the user. For example, in some embodiments, the fovea map may include a plurality of fovea regions, where each of the plurality of fovea regions includes a plurality of partitions, and each partition corresponds to a low-resolution region of the content of the corresponding partition. In some embodiments, the computing device may then input one or more of the plurality of partitions into a machine learning model that is trained to generate a super-resolution reconstruction of the fovea map based on regions of interest identified within one or more of the plurality of partitions. In one embodiment, one or more of the plurality of partitions correspond to a central fovea region of the user's approximately 30°×30° field of view (FOV).
[0026] In some embodiments, the machine learning model may include a convolutional neural network (CNN), a deep convolutional neural network (DCNN), a vision transformer (ViT), or a generative adversarial network (GAN). In some embodiments, the computing device may then output a super-resolution reconstruction of the fovea map through the machine learning model. In some embodiments, the computing device may then cause the XR display device to display one or more image frames using the super-resolution reconstruction of the fovea map. In some embodiments, the trained machine learning model may also be trained to generate a multi-frame temporal super-resolution reconstruction of the fovea map based on regions of interest identified within one or more of the plurality of partitions. For example, the multi-frame temporal super-resolution reconstruction may be applied to reconstruct the texture details of the fovea map.
[0027] As used herein, "extended reality" may refer to an electronically based form of reality that has been manipulated in some way prior to presentation to a user, which includes, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, simulated reality, immersive reality, holography, or any combination thereof. For example, "extended reality" content may include entirely computer-generated content or partially computer-generated content combined with captured content (e.g., real-world images). In some embodiments, "extended reality" content may also include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that gives a viewer a three-dimensional (3D) effect). Additionally, as used herein, it should be recognized that "extended reality" may be associated with an application, product, accessory, service, or combination thereof, which may be used, for example, to create content in extended reality and / or for use in extended reality (e.g., to perform an activity in extended reality). Thus, "extended reality" content may be implemented on various platforms, including head-mounted devices (HMDs) connected to a host computer system, standalone HMDs, mobile devices or computing systems, or any other hardware platform capable of providing extended reality content to one or more viewers.
[0028] "Passthrough" is a feature that allows a user wearing an HMD to see the characteristics of their physical environment by displaying visual information captured by the forward-facing camera of the HMD. To address misalignment between the stereo cameras and the user's eyes and provide parallax, the passthrough image is re-projected based on a 3D model representation of the physical environment. In one embodiment, the 3D model may provide rendering system geometry information, and the image captured by the HMD's camera is used as a texture image. In another embodiment, 3D information (e.g., depth information) may be utilized to apply a 2D warp to the texture. For example, the camera image may be re-projected using the 3D information of the scene to address misalignment from the camera to the user's eyes.
[0029] Figure 1AShows an example of an extended reality (XR) display device 100 worn by a user 102. In some embodiments, the extended reality (XR) display device 100 may include a head-mounted device (“HMD”) 104, a controller 106, and a computing system 108. The HMD 104 may be worn over the user's eyes and provide visual content to the user 102 through an internal display (not shown). The HMD 104 may have two separate internal displays, one for each eye of the user 102. As Figure 1A shown, the HMD 104 may completely cover the user's field of view. As the sole provider of visual information to the user 102, the HMD 104 achieves the goal of providing an immersive artificial reality experience. However, as a result, the user 102 will not be able to see the physical environment around him because his field of view is blocked by the HMD 104. Thus, the see-through feature described herein is adapted to provide the user with real-time visual information about his physical environment. The HMD 104 may include multiple outward-facing cameras 107A to 107C. In some embodiments, cameras 107A and 107B may be grayscale cameras, while camera 107C may be an RGB camera. Although multiple cameras 107 are shown, the extended reality (XR) display device 100 may include any number of cameras.
[0030] Figure 1B Shows an example of the see-through feature. The user 102 may be wearing the HMD 104 and immersed in a virtual reality environment. A physical table 150 is located in the physical environment around the user 102. However, since the user 102's field of view is blocked by the HMD 104, the user 102 cannot directly see the table 150. To help the user perceive his physical environment while wearing the HMD 104, the see-through feature uses, for example, the aforementioned outward-facing cameras 107A to 107C to collect information about the physical environment. Although the HMD 104 has three outward-facing cameras 107A to 107C, any combination of cameras 107A to 107C may be used to perform the functions described herein. For example, cameras 107A and 107B may be used to perform one or more of the functions described herein. In some embodiments, cameras 107A to 107C may be used to capture an image of the scene. Then, the collected information may be re-projected to the user 102 based on the user 102's viewpoint. For example, in some embodiments where the HMD 104 has a right display 160A for the user's right eye and a left display 160B for the user's left eye, the device 100 may respectively: (1) render a re-projected view 150A of the physical environment for the right display 160A based on the user's right-eye viewpoint; and (2) render a re-projected view 150B of the physical environment for the left display 160B based on the user's left-eye viewpoint.
[0031] Referring again to Figure 1A , the HMD 104 may have outward-facing cameras, such as Figure 1A the three forward-facing cameras 107A-107C shown. Although only three forward-facing cameras 107A-107C are shown, the HMD 104 may have any number of cameras facing in any direction (e.g., upward-facing cameras for capturing ceiling or room lighting, downward-facing cameras for capturing a portion of the user's face and / or body, rearward-facing cameras for capturing a portion of the area behind the user, and / or internal cameras for capturing the user's eye gaze direction for eye-tracking purposes). The outward-facing cameras are configured to capture the physical environment surrounding the user and may continuously make such captures to generate a series of frames (e.g., as a video). As previously explained, although the images captured by the forward-facing cameras 107A-107C may be directly displayed to the user 102 via the HMD 104, doing so may not provide the user with an accurate view of the physical environment because the cameras 107A-107C cannot physically be located in exactly the same position as the user's eyes. Thus, the see-through feature described herein uses a reprojection technique that can generate a 3D representation of the physical environment and then render an image from the viewpoint of the user's eyes based on that 3D representation.
[0032] The 3D representation may be generated based on depth measurements of physical objects observed through the cameras 107A-107C. The depth may be measured in various ways. In some embodiments, the depth may be calculated based on stereo images. For example, the three forward-facing cameras 107A-107C may share an overlapping field of view and may be configured to capture images simultaneously. Thus, the same physical object may be captured simultaneously by the cameras 107A-107C. For example, in the image captured by camera 107A, a particular feature of the object may appear at a pixel pA, while in the image captured by camera 107B, the same feature may appear at another pixel pB. As long as the depth measurement system knows that these two pixels correspond to the same feature, it may use triangulation techniques to calculate the depth of the observed feature.
[0033] For example, based on the position of the camera 107A in the 3D space and the pixel position of pA relative to the field of view of the camera 107A, a line can be projected from the camera 107A and passing through the pixel pA. A similar line can be projected from another camera 107B and passing through the pixel pB. Since it is assumed that the two pixels correspond to the same physical feature, these two lines should intersect. The two intersecting lines, together with the imaginary line drawn between the two cameras 107A and 107B, form a triangle, which can be used to calculate the distance of the observed feature from the camera 107A or 107B or a point in the space where the observed feature is located. The same operation can also be performed between the camera 107A or 107B and the camera 107C.
[0034] In some embodiments, the pose (e.g., position and orientation) of the HMD 104 within the environment may be appropriate. For example, when the user 102 moves around in the virtual environment, in order to render an appropriate display to the user 102, the device 100 may be adapted to determine the position and orientation of the user at any given moment. The device 100 may also determine the viewpoint of any one of the cameras 107A to 107C or the viewpoint of either eye of the user's binoculars based on the pose of the HMD. In some embodiments, the HMD 104 may be equipped with an inertial-measurement unit (IMU). The data generated by the IMU and the stereoscopic images captured by the outward-facing cameras 107A and 107B allow the device 100 to calculate the pose of the HMD 104 using, for example, simultaneous localization and mapping (SLAM) or other suitable techniques.
[0035] In some embodiments, the extended reality (XR) display device 100 may also have one or more controllers 106 that enable the user 102 to provide input. The controller(s) 106 may communicate with the HMD 104 or a separate computing unit 108 via a wireless connection or a wired connection. The controller(s) 106 may have any number of buttons or other mechanical input mechanisms. Additionally, the controller(s) 106 may have an IMU such that the position of the controller(s) 106 can be tracked. The controller(s) 106 may also be tracked based on a predefined pattern on the controller. For example, the controller(s) 106 may have a plurality of infrared light-emitting diodes (LEDs) or other known observable features that together form a predefined pattern. The device 100 may be able to use sensors or cameras to capture an image of the predefined pattern on the controller. The system may calculate the position and orientation of the controller relative to the sensor or camera based on the orientation of the observed patterns.
[0036] The extended reality (XR) display device 100 may also include a computer unit 108. The computer unit may be a separate unit physically separate from the HMD 104, or may be integrated with the HMD 104. In embodiments where the computer 108 is a separate unit, the computer may be communicatively coupled to the HMD 104 via a wireless link or a wired link. The computer 108 may be a high-performance device (e.g., a desktop computer or a laptop computer), or a resource-constrained device (e.g., a mobile phone). A high-performance device may have a dedicated graphics processing unit (GPU) and a high-capacity or constant power supply. On the other hand, a resource-constrained device may not have a GPU and may have a limited battery capacity. Thus, the algorithms that the extended reality (XR) display device 100 can actually use depend on the capabilities of its computer unit 108.
[0037] For example, in some embodiments where the computing unit 108 is a high-performance device, embodiments of the see-through feature may be designed as follows. A series of images of the surrounding physical environment can be captured by the outward-facing cameras 107A to 107C of the HMD 104. However, the information captured by the cameras 107A to 107C may not be aligned with the information that the user's eyes can capture, because the cameras and the user's eyes may not be spatially coincident (e.g., the cameras may be located at a certain distance from the user's eyes and thus have different viewpoints). Thus, simply displaying the content captured by the cameras to the user may not be an accurate representation of what the user should perceive.
[0038] Figure 2A and Figure 2B illustrates an example sensor readout of an image sensor according to a particular embodiment. First, refer to Figure 2A , Figure 2A illustrates an example sensor readout 200A. In some embodiments, the sensor readout 200A may include multiple partitions 202, 204, 206, 208. By way of example and not limitation, the sensor readout 200A may include a first partition 202, a second partition 204, a third partition 206, and a fourth partition 208. Although a particular number of partitions is shown, the present disclosure contemplates sensor readouts of any suitable configuration including any number of partitions. In some embodiments, partition 1 202 may include a sensor resolution 210 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 210 of partition 1 202 may be full resolution. In some embodiments, partition 2 204 may include a sensor resolution 212 that specifies a pixel pattern to be read out over an area of the image sensor.
[0039] In some embodiments, the sensor resolution 212 of partition 2 204 may be 1 / 2 resolution. In some embodiments, partition 3 206 may include a sensor resolution 214 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 214 may be 1 / 4 resolution. In some embodiments, partition 4 208 may include a sensor resolution 216 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 216 may be 1 / 8 resolution. In some embodiments, the sensor resolutions 210, 212, 214, 216 may specify one or more RGB pixels to be read out in sensor readout 200A. In some embodiments, an RGB camera may use sensor readout 200A to read out the pixels designated to generate the output image as described herein. In some embodiments, based on the respective sensor resolutions 210, 212, 214, 216, partition 1 202 may generate a first image, partition 2 204 may generate a second image, partition 3 206 may generate a third image, and partition 4 208 may generate a fourth image.
[0040] In some embodiments, an RGB camera may use sensor readout 200A to generate an output image and send the output image to the ISP for processing as described herein. In some embodiments, an image sensor using sensor readout 200A may output the active pixels designated by sensor readout 200A. In some embodiments, the position and size of a frame region of interest (ROI) may be programmable. In some embodiments, although a uniform distribution of pixels in the sensor resolutions 210, 212, 214, 216 is shown, a non-uniform distribution may also be used for sensor readout 200A. In some embodiments, an algorithm may be used, for example, by averaging, to combine multiple pixels from a full sensor readout or sensor readout 200A into a single pixel.
[0041] Reference Figure 2B , Figure 2B FIG. shows another example sensor readout 200B. In some embodiments, sensor readout 200B may include a plurality of partitions 222, 224, 226, 228. By way of example and not limitation, sensor readout 200B may include a first partition 222, a second partition 224, a third partition 226, and a fourth partition 228. Although a specific number of partitions is shown, the present disclosure contemplates sensor readouts of any suitable configuration including any number of partitions. In some embodiments, partition 1 222 may include a sensor resolution 230 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 230 of partition 1 222 may be full resolution.
[0042] In some embodiments, partition 2 224 may include a sensor resolution 232 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 232 of partition 2 224 may be 1 / 2 resolution. In some embodiments, partition 3 226 may include a sensor resolution 234 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 234 may be 1 / 4 resolution. In some embodiments, partition 4 228 may include a sensor resolution 236 that specifies a pixel pattern to be read out over an area of the image sensor. In some embodiments, the sensor resolution 236 may be 1 / 8 resolution. In some embodiments, the sensor resolutions 230, 232, 234, 236 may specify one or more monochromatic pixels to be read out in sensor readout 200B. In some embodiments, a monochrome camera may use sensor readout 200B to read out the specified pixels to generate an output image as described herein.
[0043] In some embodiments, based on the respective sensor resolutions 230, 232, 234, 236, partition 1 222 may generate a first image, partition 2 224 may generate a second image, partition 3 226 may generate a third image, and partition 4 228 may generate a fourth image. In some embodiments, a monochrome camera may use sensor readout 200B to generate an output image and send the output image to the ISP for processing as described herein. In some embodiments, an image sensor using sensor readout 200B may output active pixels specified by sensor readout 200B. In some embodiments, the position and size of the frame ROI may be programmable. In some embodiments, an algorithm may be used, such as by averaging, to combine multiple pixels from a full sensor readout or sensor readout 200B into a single pixel.
[0044] Figure 3 An example image divided into multiple different partitions according to embodiments of the present disclosure is shown. In some embodiments, image 300 may include multiple partitions 302, 304, 306. In some embodiments, image 300 may include an indication of an eye gaze representing the focus of a user's eye. In some embodiments, the eye gaze may be used to determine the fovea map described herein. The fovea map may be used to determine the sensor readout corresponding to the multiple partitions 302, 304, 306. As Figure 3As shown, image 300 has different resolutions for each partition. In partition 1 302, the resolution is 1:1. In partition 2 304, the resolution is 2:1 subsampling. In partition 3 306, the resolution is 4:1 subsampling. As shown, the farther away from the center of the user's eye gaze, the lower the resolution of the output image. In some embodiments, image 300 may represent the output image from a camera. Image 300 may be processed by a computing system to generate an output display image presented to a user of the computing system as described herein.
[0045] Figure 4A and Figure 4B illustrates an example process for foveal region processing according to embodiments of the present disclosure. Referring to Figure 4A , process 400 for foveal region processing may be performed by the computing system described herein. In some embodiments, process 400 may include camera image processing 402, depth processing 404, view estimated from eye position 406 (as Figure 4B shown), and display 410 (as Figure 4B shown). In some embodiments, process 400 may perform camera image processing 402 and depth processing 404 simultaneously or in any particular order. In some embodiments, process 400 may start by acquiring a monochrome image from monochrome camera 412 and an RGB image from color camera 414. The images acquired by the monochrome camera and color camera 414 may be sent to the ISP for processing 402.
[0046] In some embodiments, the computing system may perform eye tracking and determine the user's gaze 420, which is used for foveal acquisition 416 corresponding to the monochrome image acquired by monochrome camera 412 and foveal acquisition 418 corresponding to the RGB image acquired by color camera 414. Foveal acquisition 416 may generate a foveal map based on the determined user's gaze 420 and use the generated foveal map to determine sensor readout for monochrome camera 412 as described herein. The output of foveal acquisition 416 may be a monochrome image with reduced resolution. Foveal acquisition 418 may generate a foveal map based on the determined user's gaze 420 and use the generated foveal map to determine sensor readout for color camera 414 as described herein. The output of foveal acquisition 418 may be an RGB image with reduced resolution. Foveal acquisitions 416, 418 may output to corresponding temporal noise reduction (TNR) 422, 424. TNR 422, 424 may receive depth and lens distortion 426 to perform TNR. The outputs of TNR 422, 424 may be combined in chrominance and luminance fusion 428. Chrominance and luminance fusion 428 may receive depth and distortion 430 to perform chrominance and luminance fusion 428.
[0047] In some embodiments, the depth and lens distortion 426 can be the same as the depth and distortion 430. In some embodiments, the chromaticity and luminance fusion 428 can output a high-resolution, foveated, denoised RGB image in the camera space 432 for processing by the warping process 406. Although only one set of monochromatic cameras 412 and color cameras 414 are shown in this process 400, additional sets of monochromatic cameras 412 and color cameras 414 can be present for acquiring images for use in the camera image processing 402. By way of example and not limitation, one set of cameras can be used for each eye.
[0048] In some embodiments, an indirect time-of-flight (iTOF) camera 434 can generate distance measurements for use in the depth processing 404. Although this disclosure describes the use of iTOF, this disclosure contemplates the use of direct time-of-flight (dTOF) or any time-of-flight technology. In some embodiments, a stereo camera 436 can acquire stereo images for use in the depth processing 404. In some embodiments, the stereo images can be sent to the foveated stereo process 438. The foveated stereo process 438 can receive the determined user's fixation 440 (which can be the same as the determined user's fixation 420). As described herein, the foveated stereo process 438 can generate a foveated map to determine the sensor readout of each of the plurality of stereo cameras 436. The outputs of the iTOF camera 434 and the foveated stereo process 438 are sent to the machine learning densification process 442.
[0049] The output of the machine learning (ML) densification process 442 is sent to the optical flow process 444 and the ML segmentation process 446. The outputs of the optical flow process 444 and the ML segmentation process 446 are combined using the time domain and stabilization 448 to generate a temporally stable depth map 450. The temporally stable depth map 450 undergoes 3D warping to the camera space process 454. The output of the 3D warping to the camera space process 454 undergoes an upsampling and densification process 452. The output of the upsampling and densification process 452 is used for the depth and lens distortion 426, 430. The output of the upsampling and densification process 452 is combined with the high-resolution, foveated, denoised RGB image 432 in the camera space by the 3D warping to the eye space process 456. The 3D warping to the eye space process 456 can receive the rendering pose 458. The 3D warping to the eye space process 456 warps the image to the baseline of the user's eye. The output of the 3D warping to the eye space process 456 can be sent to the ML inpainting process 460. The ML inpainting process 460 corrects the warping artifacts that arise from the warping process 406.
[0050] In some embodiments, the output of the ML repair process 460 can be sent to the rendering process 408 as shown Figure 4B The output of the ML repair process 460 can undergo a compositing process 462 in which one or more virtual objects are rendered within the image. Then, the composite image from the compositing process 462 can be placed in a super-resolution process 464 in which one or more super-resolution techniques can be applied to the composite image. The super-resolution process 464 can use the user's current updated gaze 466 to perform one or more super-resolution techniques.
[0051] For example, in one embodiment, the super-resolution process 464 can apply different super-resolution techniques to the composite image to equalize the resolution of the entire image. In another embodiment, independent of whether visual foveation is used on cameras 412, 414, the super-resolution process 464 can also include increasing the resolution beyond the input level and / or maintaining the same resolution but enhancing the high spatial frequencies present in the image. For example, if the resolution is different due to foveal acquisition, the super-resolution process 464 applies one or more super-resolution techniques to increase the resolution of the composite image. The output of the super-resolution process 464 is sent to the lens distortion and head rotation correction process 468 which performs lens distortion and head rotation correction based on the user's head pose and lens distortion. In some embodiments, in the display process 410, the image is displayed on one or more displays of the computing system. In some embodiments, process 400 can be performed on the cameras for each eye of the user to render an output image corresponding to the passthrough function.
[0052] In some embodiments, one or more of the foregoing techniques may include applying a super-resolution process 464 to the synthetic image as a whole or selectively. Specifically, as generally discussed above, super-resolution may be applied to the synthetic image. Each of the following embodiments has advantages. For example, in one embodiment, a synthetic image may be generated, and then super-resolution may be applied to the combined image (regardless of layers). In another embodiment, a synthetic image may be generated, and then super-resolution may be applied to the combined image, but selectively to frames and / or regions having certain layers (such as pass-through). In one example, one solution may be to add a super-resolution block for each layer before generating the synthetic image. Applying super-resolution after generating the synthetic image may have the advantage of applying super-resolution at a single location, but it also applies to all content of the generated synthetic image. Similarly, generating a synthetic image and then selectively applying super-resolution to frames and / or regions has a lower latency (e.g., from time warping to display), and may have specialized algorithms based on layer type. Specifically, applying the super-resolution process 464 to all or selectively to selected image frames may consume a large amount of computing resources and cause the XR display device 100 to consume relatively more power. According to the presently disclosed embodiments, it may be useful to provide gaze-based super-resolution for the image content rendered for the XR display device 100.
[0053] Figure 5 FIG. shows a gaze-based super-resolution reconstruction model 500 according to presently disclosed embodiments, which may be applicable to provide gaze-based super-resolution for image content rendered by an extended reality (XR) device. In certain embodiments, the gaze-based super-resolution reconstruction model 500 may include a low-resolution RGB image 502 and a low-resolution monochrome image 504 input into a super-resolution reconstruction model 506. In certain embodiments, as Figure 5 shown by the gaze-based super-resolution reconstruction model 500 of, the low-resolution RGB image 502 and the low-resolution monochrome image 504 may be input into a machine learning model 506 and preprocessed by the machine learning model 506.
[0054] In certain embodiments, preprocessing the low-resolution RGB image 502 and the low-resolution monochrome image 504 may include normalizing and scaling the depth values and color values of the low-resolution RGB image 502 and the low-resolution monochrome image 504, respectively. In certain embodiments, as Figure 5As further shown in the gaze-based super-resolution reconstruction model 500, the low-resolution RGB image 502 and the low-resolution monochrome image 504 can be concatenated (e.g., stacked) and input into the machine learning model 506. For example, in some embodiments, concatenating (e.g., stacking) the low-resolution RGB image 502 and the low-resolution monochrome image 504 and inputting the low-resolution RGB image 502 and the low-resolution monochrome image 504 into the machine learning model 506 simultaneously can reduce the model complexity and execution time.
[0055] In certain embodiments, the machine learning model 506 can include a convolutional neural network (CNN), a deep convolutional neural network (DCNN), a vision transformer (ViT), a generative adversarial network (GAN), or other similar deep neural networks, and these other similar deep neural networks can be applicable to encode the low-resolution RGB image 502 and the low-resolution monochrome image 504 into one or more feature maps or a series of image patches, and the one or more feature maps or the series of image patches can be classified and used to generate predictions of the high-resolution output image 520. For example, the machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.) can be trained to generate estimates of the reconstructed high-resolution RGB image and monochrome image, and these reconstructed high-resolution RGB image and monochrome image include both the high-resolution monochrome information and the RGB color information of a specific region of interest. Therefore, according to the currently disclosed technology, the machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.) can be trained to reconstruct the low-resolution RGB image and monochrome image into an RGB image and a monochrome image including the high-resolution region of interest based on the eye gaze of the user 102.
[0056] For example, as described above, based on the low-resolution RGB image 502 and the low-resolution monochrome image 504, the machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.) can include one or more downsampling layers 508, and the one or more downsampling layers 508 can generate one or more feature maps or a series of image patches based on the low-resolution RGB image 502 and the low-resolution monochrome image 504. Specifically, as Figure 5 As further shown in the gaze-based super-resolution reconstruction model 500, the downsampling layer 508 can generate one or more sets of interpolation weights 510, 512, and 514 as its predictions for each pixel in the low-resolution RGB image 502 and the low-resolution monochrome image 504 corresponding to the high-resolution output image 520. For example, in one embodiment, if the high-resolution output image 520 has N×N pixels, the interpolation weights 510, 512, and 514 can represent 1000×1000 sets of interpolation weights (e.g., each set can have N interpolation parameters).
[0057] As previously described, the machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.) will ultimately generate a prediction of the high-resolution reconstruction of the region of interest corresponding to the low-resolution RGB image 502 and the low-resolution monochrome image 504, which corresponds to the region pointed to by the eye gaze of the user 102 or predicted to be pointed to within the next several or more image frames. In some embodiments, as Figure 5 further shown by the gaze-based super-resolution reconstruction model 500 of, one or more sets of interpolation weights 510, 512, and 514 can then be provided to one or more upsampling layers 516 for upsampling, and the output values can be input into the interpolation function block 518. In some embodiments, the interpolation function block 518 can be any function block suitable for algorithmically calculating the color of each pixel of the high-resolution output image 520 by interpolating the downsampled output values. In one embodiment, the interpolation function block 518 can algorithmically calculate the color of each pixel of the region of interest of the high-resolution output image 520.
[0058] In this way, the current embodiment can provide gaze-based super-resolution for the image content rendered by the XR display device 100. Specifically, by tracking the eye gaze of the user 102 and the image content being displayed by the XR display device 100 as a trigger for when to apply the gaze-based super-resolution reconstruction techniques described herein, gaze-based super-resolution reconstruction can be applied to the central fovea region of the 30°×30° field of view (FOV), so it can be efficiently executed and reduce the power consumption of the XR display device 100. That is, the currently disclosed gaze-based super-resolution reconstruction technique can be efficiently executed and reduce the power consumption of the XR display device 100 while always rendering high-resolution images to the user 102. Therefore, the current embodiment can apply different pixel scaling or image reconstruction techniques based on the content in the scene.
[0059] In some embodiments, the machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.) can be further trained to reconstruct low-resolution RGB images and monochromatic images into RGB images and monochromatic images including high-resolution regions of interest based on the pose of an object, e.g., relative to the image content being rendered and displayed by the XR display device 100, in addition to being trained to reconstruct low-resolution RGB images and monochromatic images into RGB images and monochromatic images including high-resolution regions of interest based on the eye gaze of the user 102. For example, in some embodiments, the machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.) can identify regions of interest that include visually interesting or visually prominent image content that the user 102 may gaze at. For example, current gaze-based super-resolution reconstruction techniques can be applied to image content such as people, animals, the faces of people or animals, the eyes of people or animals, devices, multicolored flowers, text, and logos. Thus, the current embodiments can apply different pixel scaling or image reconstruction techniques based on the eye gaze of the user 102.
[0060] In some embodiments, the upsampling according to the currently disclosed embodiments can be based on previously acquired key frames for inpainting or be trained on a pre-acquired environmental model. Specifically, the MR pipeline for rendering the see-through image can be adjusted to address occlusion issues. Due to the difference between the position of the see-through camera (e.g., the camera for acquiring the image that will be distorted to generate the see-through image) and the eyes of the user 102, once the acquired images are distorted to the viewpoint of the eyes of the user 102, these acquired images may lose information. Even if these cameras are placed as close as possible to the eyes of the user 102, there will inevitably be a difference. For example, when the scene includes a foreground object (e.g., the hand of the user 102), the amount of background occluded by the foreground object will be significantly different from the perspective of the see-through camera and the eyes of the user 102. Therefore, when the images acquired by the see-through camera are re-projected to the eye position of the user 102, some parts of the background occluded by the foreground object in the acquired images will become visible to the user 102, resulting in the aforementioned occlusion problem.
[0061] In some embodiments, the MR pipeline can address occlusion problems by leveraging scene information collected in previous frames. At a high level, the idea is to use previously captured images to provide information about the background scene that may be occluded in the current image capture. When user 102 dons the head-mounted viewer or starts a mixed reality session, the external camera will capture an image of user 102's scene. The 3D reconstruction module can scan the environment of user 102, collecting image data and depth measurements. In some embodiments, the images and depth data can be filtered to separate the background from foreground objects. Then, the image and depth data associated with the background can be used to generate a 3D model (e.g., a mesh) of the background and a corresponding texture atlas. Using the 3D model and texture atlas, the rendering system will be able to synthesize an image of the background scene from a new viewpoint. Then, the pixel information of the background in the synthesized image can be used to fill in the missing pixel information in the passthrough image.
[0062] In some embodiments, the MR pipeline can use previously captured red, green, blue, and depth (RGBD) key frames for inpainting instead of using past observations to build a 3D mesh model of the background environment. The camera and depth sensor of the head-mounted viewer can capture red, green, blue, and depth (RGBD) information of the scene. The key frame selection logic module can process the RGBD images to select which images are suitable for inpainting the background content of subsequent frames. The selected RGBD images (which can be referred to as key frames) are saved to a collection over time. Similar to the 3D mesh reconstruction of the background, the collection of key frames represents the historical observations of the background from different viewpoints. Thus, ideally, each selected key frame contains only the background and no foreground objects. Therefore, the task of the key frame selection logic module can be to analyze each input RGBD image and find those that contain only background content. However, in some use cases, it may be difficult to find RGBD images without foreground objects. Thus, in certain embodiments, the criteria for selecting key frames can be relaxed to also include RGBD images that mainly include background content. The MR pipeline can use the collection of key frames to render the background information of the passthrough image from the perspective of user 102's eyes. The MR pipeline can use any suitable rendering technique, such as point cloud-based rendering, neural volume rendering, unstructured Lumigraph rendering, etc., to render the background image using the collection of key frames.
[0063] The super-resolution process can utilize past observations of the background, which can be represented using the aforementioned 3D reconstruction model or a set of key frames, to improve the resolution and / or quality of the passthrough image. For example, when the MR pipeline determines that the gaze of user 102 is focused on the background, the MR pipeline can improve the resolution of the region of interest by using the 3D reconstruction model or key frames to render the corresponding pixel values. The MR pipeline can fuse the rendered pixel values with the final passthrough image in a variety of ways. For example, the rendered pixel values of the background and the low-resolution image acquisition can be jointly provided to a machine learning model 506 (e.g., CNN, DCNN, ViT, and GAN, etc.), which is trained to consider these two types of information when generating the super-resolution output image 520. As another example, the rendered pixel values of the background can be directly used as the background of the passthrough image and / or the synthetic image. In yet another example, the rendered pixel values of the background can be mixed with the existing background of the passthrough image and / or the synthetic image.
[0064] Figure 6A and Figure 6B Respectively illustrate one or more image reconstruction examples 600A and 600B according to the presently disclosed embodiments. Specifically, the presently disclosed gaze-based super-resolution reconstruction technique can be compared with existing pixel scaling or image reconstruction examples. For example, image 602 can include a low-resolution original image to be magnified. In one embodiment, image 602 can be magnified according to the nearest neighbor interpolation technique, resulting in image 604 (e.g., a blurry image). In another embodiment, image 602 can be magnified according to a general super-resolution reconstruction technique, resulting in image 606. Although image 606 may have a higher resolution and image quality compared to image 604, applying the general super-resolution reconstruction technique to all or part of image 602 may be computationally costly and cause the XR display device 100 to consume relatively more power. Therefore, in certain embodiments, the presently disclosed gaze-based super-resolution reconstruction technique can be applied to image 602, resulting in image 608. As shown, the presently disclosed gaze-based super-resolution reconstruction technique generates image 608, which includes high-resolution pixels in the region of interest of image 602 (e.g., the region pointed to by the eye gaze of user 102) and low-resolution pixels in the remaining region. Therefore, the presently disclosed gaze-based super-resolution reconstruction technique can be efficiently executed and reduce the power consumption of the XR display device 100 while always rendering a high-resolution image to user 102.
[0065] Figure 7A flowchart of a method 700 for providing gaze-based super resolution for image content rendered for an extended reality (XR) device according to embodiments of the present disclosure is shown. Method 700 may be executed using one or more processing devices (e.g., XR display device 100), which may include hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), or any one or more other processing devices suitable for processing image data), software (e.g., instructions running / executing on one or more processors), firmware (e.g., microcode), or some combination thereof.
[0066] Method 700 may begin at block 702, where one or more processing devices (e.g., XR display device 100) render an extended reality (XR) environment on one or more displays of a computing device. Then, method 700 may continue at block 704, where one or more processing devices (e.g., XR display device 100) determine the context of the XR environment relative to the user. For example, in some embodiments, determining the context may include determining one or more characteristics associated with at least one eye of the user relative to the content of reality within the XR environment. Then, method 700 may continue at block 706, where one or more processing devices (e.g., XR display device 100) generate a foveal map based on one or more characteristics associated with at least one eye of the user. For example, in one embodiment, the foveal map may include a plurality of foveal regions, which may include, for example, a plurality of partitions, each corresponding to a low-resolution region of the content of the respective partition.
[0067] The method 700 may then continue at block 708, where one or more processing devices (e.g., the XR display device 100) input one or more of the plurality of partitions into a machine learning model that is trained to generate a super-resolution reconstruction of the fovea map based on regions of interest identified within the one or more of the plurality of partitions. The method 700 may then end at block 710, where one or more processing devices (e.g., the XR display device 100) outputs a super-resolution reconstruction of the fovea map via the machine learning model. For example, the XR display device 100 may then display one or more image frames using the super-resolution reconstruction of the fovea map.
[0068] Figure 8 An example computer system 800 is shown that can be used to perform one or more of the various aforementioned techniques as currently disclosed herein. In some embodiments, one or more computer systems 800 perform one or more steps of one or more methods described or illustrated herein. In some embodiments, one or more computer systems 800 provide functionality described or illustrated herein. In some embodiments, software running on one or more computer systems 800 performs one or more steps of one or more methods described or illustrated herein, or provides functionality described or illustrated herein. Certain embodiments include one or more portions of one or more computer systems 800. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Furthermore, reference to a computer system may encompass one or more computer systems, where appropriate.
[0069] The present disclosure contemplates any suitable number of computer systems 800. The present disclosure contemplates computer systems 800 in any suitable physical form. By way of example and not limitation, computer system 800 can be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or a system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these systems. In appropriate instances, computer system 800: can include one or more computer systems 800; can be single or distributed; can span multiple locations; can span multiple machines; can span multiple data centers; or can be located in a cloud, which can include one or more cloud components in one or more networks. In appropriate instances, one or more computer systems 800 can perform one or more steps of one or more of the methods described or shown herein without substantial spatial or temporal limitations.
[0070] By way of example and not limitation, one or more computer systems 800 can perform one or more steps of one or more of the methods described or shown herein in real time or in a batch processing mode. In appropriate instances, one or more computer systems 800 can perform one or more steps of one or more of the methods described or shown herein at different times or at different locations. In certain embodiments, computer system 800 includes a processor 802, a memory 804, a storage device 806, an input / output (I / O) interface 808, a communication interface 810, and a bus 812. Although the present disclosure describes and shows a particular computer system having a particular number of particular components in a particular arrangement, the present disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0071] In some embodiments, the processor 802 includes hardware for executing instructions (e.g., those that make up a computer program). By way of example and not limitation, to execute instructions, the processor 802 may retrieve (or fetch) instructions from internal registers, internal caches, memory 804, or storage device 806; decode and execute those instructions; and then write one or more results to internal registers, internal caches, memory 804, or storage device 806. In some embodiments, the processor 802 may include one or more internal caches for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates a processor 802 that includes any suitable number of any suitable internal caches. By way of example and not limitation, the processor 802 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The instructions in the instruction cache may be copies of the instructions in memory 804 or storage device 806, and the instruction cache may accelerate the retrieval of these instructions by the processor 802.
[0072] The data in the data cache may be a copy of the data in memory 804 or storage device 806 for use by instructions executed at the processor 802; may be the result of a previous instruction executed at the processor 802 for access by a subsequent instruction executed at the processor 802 or for writing to memory 804 or storage device 806; or may be other suitable data. The data cache may accelerate read or write operations of the processor 802. The TLB may accelerate virtual address translation of the processor 802. In some embodiments, the processor 802 may include one or more internal registers for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates a processor 802 that includes any suitable number of any suitable internal registers. In appropriate instances, the processor 802 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include more than one processor 802. Although the present disclosure describes and illustrates particular processors, the present disclosure contemplates any suitable processor.
[0073] In some embodiments, the memory 804 includes a main memory that stores instructions for execution by the processor 802 or data for operation by the processor 802. By way of example and not limitation, the computer system 800 can load instructions from the storage device 806 or another source (e.g., another computer system 800) into the memory 804. Then, the processor 802 can load these instructions from the memory 804 into internal registers or an internal cache. To execute these instructions, the processor 802 can retrieve and decode the instructions from the internal registers or the internal cache. During or after execution of these instructions, the processor 802 can write one or more results (which can be intermediate or final results) to the internal registers or the internal cache. Then, the processor 802 can write one or more of these results to the memory 804. In some embodiments, the processor 802 executes only instructions in one or more of the internal registers or the internal cache, or in the memory 804 (and not in the storage device 806 or elsewhere), and operates only on data in one or more of the internal registers or the internal cache, or in the memory 804 (and not in the storage device 806 or elsewhere).
[0074] One or more memory buses (which can each include an address bus and a data bus) can couple the processor 802 to the memory 804. As described below, the bus 812 can include one or more memory buses. In some embodiments, one or more memory management units (MMUs) are located between the processor 802 and the memory 804 and facilitate access to the memory 804 requested by the processor 802. In some embodiments, the memory 804 includes random access memory (RAM). In appropriate circumstances, the RAM can be volatile memory. In appropriate circumstances, the RAM can be dynamic RAM (DRAM) or static RAM (SRAM). Additionally, in appropriate circumstances, the RAM can be single-port RAM or multi-port RAM. The present disclosure contemplates any suitable RAM. In appropriate circumstances, the memory 804 can include one or more memories 804. Although the present disclosure describes and illustrates particular memories, the present disclosure contemplates any suitable memory.
[0075] In some embodiments, storage device 806 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 806 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these storage devices. In appropriate instances, storage device 806 can include removable media or non-removable (or fixed) media. In appropriate instances, storage device 806 can be internal or external to computer system 800. In some embodiments, storage device 806 is non-volatile solid-state memory. In some embodiments, storage device 806 includes read-only memory (ROM). In appropriate instances, the ROM can be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these ROMs. The present disclosure contemplates a mass storage device 806 in any suitable physical form. In appropriate instances, storage device 806 can include one or more storage control units that facilitate communication between processor 802 and storage device 806. In appropriate instances, storage device 806 can include one or more storage devices 806. Although the present disclosure describes and shows particular storage devices, the present disclosure contemplates any suitable storage device.
[0076] In some embodiments, I / O interface 808 includes hardware, software, or both that provide one or more interfaces for communication between computer system 800 and one or more I / O devices. In appropriate instances, computer system 800 may include one or more of these I / O devices. One or more of these I / O devices may enable communication between a person and computer system 800. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more of these I / O devices. I / O devices may include one or more sensors. The present disclosure contemplates any suitable I / O devices and any suitable I / O interface 808 for these I / O devices. In appropriate instances, I / O interface 808 may include one or more device or software drivers that enable processor 802 to drive one or more of these I / O devices. In appropriate instances, I / O interface 808 may include one or more I / O interfaces 808. Although the present disclosure describes and shows particular I / O interfaces, the present disclosure contemplates any suitable I / O interface.
[0077] In some embodiments, communication interface 810 includes hardware, software, or both that provide one or more interfaces for communication (e.g., packet-based communication) between computer system 800 and one or more other computer systems 800 or one or more networks. By way of example and not limitation, communication interface 810 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network (e.g., a Wi-Fi network). The present disclosure contemplates any suitable network and any suitable communication interface 810 for that network.
[0078] By way of example and not limitation, computer system 800 may communicate with the following networks: ad hoc networks, personal area networks (PANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), or one or more portions of the Internet, or combinations of two or more of these networks. One or more portions of one or more of these networks may be wired or wireless. For example, computer system 800 may communicate with the following networks: wireless PANs (WPANs) (e.g., Bluetooth WPANs), Wi-Fi networks, WiMAX networks, cellular telephone networks (e.g., Global System for Mobile Communications (GSM) networks), or other suitable wireless networks, or combinations of two or more of these networks. In appropriate instances, computer system 800 may include any suitable communication interface 810 for any of these networks. In appropriate instances, communication interface 810 may include one or more communication interfaces 810. Although the present disclosure describes and shows particular communication interfaces, the present disclosure contemplates any suitable communication interface.
[0079] In some embodiments, bus 812 includes hardware, software, or both that couple the various components of computer system 800 to each other. By way of example and not limitation, bus 812 may include: Accelerated Graphics Port (AGP) or other graphics bus, Enhanced Industry Standard Architecture (EISA) bus, front-side bus (FSB), HYPERTRANSPORT (HT) interconnect, Industry Standard Architecture (ISA) bus, INFINIBAND interconnect, low-pin-count (LPC) bus, memory bus, MicroChannel Architecture (MCA) bus, Peripheral Component Interconnect (PCI) bus, high-speed PCI (PCI-Express, PCIe) bus, serial advanced technology attachment (SATA) bus, Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these buses. Where appropriate, bus 812 may include one or more buses 812. Although the present disclosure describes and shows specific buses, the present disclosure contemplates any suitable bus or interconnect.
[0080] In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based integrated circuits (ICs) or other ICs (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or secure digital drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, the computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0081] In this document, unless otherwise expressly stated or the context otherwise indicates, "or" is inclusive rather than exclusive. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A or B" means "A, B, or both". Further, unless otherwise expressly stated or the context otherwise indicates, "and" is both joint and several. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A and B" means "A and B, jointly or severally".
[0082] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments described or shown herein, which would be understood by those skilled in the art. The scope of the present disclosure is not limited to the example embodiments described or shown herein. Additionally, although the present disclosure describes and shows the various embodiments herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any of the components, elements, features, functions, operations, or steps described or shown anywhere herein that would be understood by a person of ordinary skill in the art. Further, the recitation in the appended claims of a device or system, or a component in a device or system, that is adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function includes that device, system, component, whether or not the particular function is activated, turned on, or unlocked, so long as the device, system, or component is so adapted to, arranged to, capable of, configured to, enabled to, operable, or operative. Additionally, although the present disclosure describes or shows certain embodiments as providing particular advantages, certain embodiments may not provide these advantages, or may provide some or all of these advantages.
Claims
1. A method, comprising: Via a computing device configured to be worn by a user: Render an extended reality (XR) environment on one or more displays of the computing device; Determine the context of the XR environment relative to the user, wherein determining the context includes determining one or more characteristics associated with at least one eye of the user relative to the content displayed within the XR environment; Generate a fovea map based on the one or more characteristics associated with at least one eye of the user, wherein the fovea map includes a plurality of fovea regions, and wherein the plurality of fovea regions includes a plurality of partitions, each partition corresponding to a low-resolution region of the content of the respective partition; Input one or more of the plurality of partitions into a machine learning model, the machine learning model being trained to generate a super-resolution reconstruction of the fovea map based on regions of interest identified within one or more of the plurality of partitions; and Output, via the machine learning model, the super-resolution reconstruction of the fovea map.
2. The method according to claim 1, further comprising: Causing an XR display device to display one or more image frames using the super-resolution reconstruction of the fovea map.
3. The method according to claim 1 or 2 further comprises: Determine the context of the XR environment relative to the user based on a prediction of the position of at least one eye of the user relative to the content over a plurality of image frames to be displayed next.
4. The method according to any one of the preceding claims, wherein, Determining the context of the XR environment relative to the user includes determining one or more of the following: the head pose of the user; the eye gaze of the user; or the depth of eye vector convergence.
5. The method according to any one of the preceding claims, wherein, The machine learning model is further trained to generate a multi-frame temporal super-resolution reconstruction of the fovea map based on the regions of interest identified within one or more of the plurality of partitions, the multi-frame temporal super-resolution reconstruction being applied to reconstruct the texture details of the fovea map.
6. The method according to any one of the preceding claims, wherein, One or more of the plurality of partitions correspond to a central fovea region of the user's approximately 30°×30° field of view (FOV).
7. The method according to any one of the preceding claims, wherein, The machine learning model includes a convolutional neural network (CNN), a deep convolutional neural network (DCNN), a vision transformer (ViT), or a generative adversarial network (GAN).
8. A computing device configured to be worn by a user, comprising: One or more displays; One or more non-transitory computer-readable storage media, the one or more non-transitory computer-readable storage media including instructions; And One or more processors, the one or more processors coupled to the one or more displays, the one or more processors configured to execute the instructions to perform the method of any of the preceding claims.
9. A computer-readable medium, the computer-readable medium including instructions that, when executed by one or more processors of a computing device configured to be worn by a user, cause the computing device to perform the method according to any one of claims 1 to 7.
10. A computer program product comprising instructions that, when executed by one or more processors of a computing device configured to be worn by a user, cause the computing device to perform the method according to any one of claims 1 to 7.