Method and system for performing image reprojection using head posture information

By employing depth-based coding with head pose information, augmented reality systems effectively reproject images to align with user head posture, addressing the challenge of integrating virtual content with real-world elements and enhancing user experience.

JP2026513028APending Publication Date: 2026-04-22MAGIC LEAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2024-03-19
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing augmented reality systems struggle to seamlessly integrate virtual content with real-world elements due to the complexity of the human visual system, leading to discomfort and unnatural presentation of virtual image elements.

Method used

The integration of depth-based coding/decoding processes, such as Multiview High Efficiency Video Coding (MV-HEVC), combined with head pose information, allows for image reprojection that adjusts to the user's current head posture, reducing dependence on graphics processing units (GPUs) and minimizing power and area requirements.

Benefits of technology

This approach enhances the user experience by accurately reprojecting images based on head posture, improving the integration of virtual content with real-world surroundings, thereby reducing computational and power demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026513028000001_ABST
    Figure 2026513028000001_ABST
Patent Text Reader

Abstract

The method includes, in an encoder, receiving virtual content; in an encoder, receiving a predicted head pose corresponding to the virtual content; and encoding the virtual content based on the predicted head pose. The method also includes, generating compressed content; in a decoder, receiving the compressed content; and in a decoder, receiving the current head pose. The method also includes, decoding the compressed content based on the current head pose; and generating reprojected virtual content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 454,594, titled "Method and System for Performing Image Reprojection Using Head pose Information", filed on March 24, 2023, and U.S. Provisional Patent Application No. 63 / 453,376, titled "Method and System for Performing Foveated Image Compression Based on Eye Gaze", filed on March 20, 2023, the disclosures of which are hereby incorporated by reference in their entirety for all purposes.

Background Art

[0002]

[0002] Modern computing and display technologies have facilitated the development of systems for so - called virtual reality or augmented reality experiences, where digitally reproduced images or portions thereof are presented to viewers in a way that they appear or are perceived as if they were real. Virtual reality, i.e., VR scenarios, typically involve presenting digital or virtual image information without transparency to other actual real - world visual inputs, and augmented reality, i.e., AR scenarios, typically involve presenting digital or virtual image information as an augmentation to the visualization of the actual world around the viewer.

[0003]

[0003] Referring to Figure 1, an augmented reality scene 100 is depicted. The user of the AR technology sees a setting like a real-world park, featuring people, trees, buildings, and a concrete platform 120 in the background. The user also perceives that they are "seeing" "virtual content," such as a robot statue 110 standing on the real-world concrete platform 120, and a flying cartoonish avatar character 102 that appears to be an anthropomorphic bumblebee. These elements 110 and 102 are "virtual" in the sense that they do not exist in the real world. Because the human visual system is complex, it is difficult to produce AR technology that facilitates a rich presentation of virtual image elements surrounded by other virtual or real-world image elements in a comfortable and natural sense.

[0004]

[0004] Despite these advances in display technology, there is a need in the art for improved methods and systems related to augmented reality systems, particularly display systems. [Overview of the project]

[0005]

[0005] The present invention generally relates to methods and systems relating to projection display systems, including wearable displays. More specifically, embodiments of the present invention provide methods and systems useful for reprojecting images or videos from a viewpoint corresponding to the current head posture. The present invention is applicable to a variety of applications in computer vision and image display systems.

[0006]

[0006] Some embodiments of the present invention utilize a depth-based coding / decoding process, such as the Multiview High Efficiency Video Coding (MV-HEVC) process, to adjust image reprojection based on head pose. An example of the MV-HEVC coding / decoding standard is MPEG-H Part 2, also known as H.265 and MV-HEVC Part H, which was implemented as part of the MPEG-H project as a successor to Advanced Video Coding (AVC, H.264, or MPEG-4 Part 10). The MPEG MV-HEVC standard uses depth-based reprojection to render images from both left and right viewpoints and enables compression of 3D recorded video.

[0007]

[0007] Embodiments of the present invention combine the existing MPEG H.265 standard for depth map reprojection with head pose information. Adding head pose information with depth map reprojection makes it possible to perform head pose-correlated image reprojection using existing hardware. By utilizing MV-HEVC hardware, for example, MPEG H.265 hardware or stage modified to utilize head pose information, embodiments of the present invention can implement image reprojection, including head pose-correlated 6DOF final stage image reprojection correction, with reduced or no dependence on graphics processing units (GPUs) or other conventional processors. By reducing or eliminating GPU dependence, embodiments of the present invention reduce both area and power requirements compared to conventional systems. Although the embodiments described herein are described in the context of H.265, this is not essential, and other suitable encoding / decoding systems, including depth-based compression systems and other encoding / decoding systems that utilize multiple views, can be used within the scope of the present invention.

[0008]

[0008] Accordingly, the embodiments described herein use a depth-based MPEG reprojection process combined with head posture adjustment. As a result, the video reprojection takes the current head posture into account.

[0009]

[0009] The present invention achieves many advantages over conventional techniques. For example, embodiments of the present invention provide a method and system that enables the reprojection of an image or video based on the user's current head posture, thereby improving the user experience. These and other embodiments of the present invention, along with many of their advantages and features, will be described in more detail below in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0010] [Figure 1] This diagram shows a user's view of augmented reality (AR) through an AR device. [Figure 2A] These are cross-sectional side views of an example of a stacked set of waveguides, each containing an internally coupled optical element. [Figure 2B] Figure 2A is a perspective view of an example of one or more stacked waveguides. [Figure 2C] Figures 2A and 2B are top views of an example of one or more stacked waveguides. [Figure 3] This is a simplified diagram of an eyepiece waveguide with combined pupil expanders according to one embodiment of the present invention. [Figure 4] This figure shows an example of a wearable display system according to one embodiment of the present invention. [Figure 5] This is a perspective view of a wearable device according to one embodiment of the present invention. [Figure 6] This is a simplified diagram illustrating the reprojection of virtual content using one embodiment of the present invention. [Figure 7] This is a simplified schematic diagram showing an image reprojection system according to one embodiment of the present invention. [Figure 8] This is a simplified flowchart illustrating a method for encoding and decoding an image based on head pose, according to one embodiment of the present invention. [Figure 9] This is a simplified schematic diagram illustrating the operation of an image reprojection system according to one embodiment of the present invention. [Figure 10]This is a diagram showing images compressed using a single quality setting. [Figure 11] This diagram shows a fovealed image having three fovealed regions according to one embodiment of the present invention. [Figure 12] This is a fovealed 3D generated image having three fovealed regions, according to yet another embodiment of the present invention. [Figure 13] This is a diagram showing an image that can be used with multiple foveal maps according to one embodiment of the present invention. [Figure 14] This is a simplified flowchart illustrating a method for compressing images using one embodiment of the present invention. [Figure 15] This figure shows the compression level obtained as a function of time, expressed by consecutive frames versus frequency, for both a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. [Figure 16] The following shows a histogram of frame count versus compression for a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. [Figure 17] This is a simplified flowchart illustrating a method for compressing image frames using an alternating compression algorithm according to one embodiment of the present invention. [Figure 18] This is a simplified image showing an image frame divided into high-quality and low-quality regions according to one embodiment of the present invention. [Figure 19] This is a simplified flowchart illustrating a method for compressing an image using different compression ratios for high-quality and low-quality regions, according to one embodiment of the present invention. [Figure 20] This is a simplified image showing an image frame divided into high-quality tiles and low-quality tiles according to one embodiment of the present invention. [Figure 21] This is a simplified flowchart illustrating a method for compressing images using different compression ratios for high-quality and low-quality tiles according to one embodiment of the present invention. [Figure 22]A series of images showing the use of a spherical depth map according to an embodiment of the present invention. [Figure 23] A simplified schematic diagram showing the components of an AR system according to an embodiment of the present invention.

Mode for Carrying Out the Invention

[0011]

[0035] The present invention generally relates to methods and systems related to projection display systems including wearable displays. More particularly, embodiments of the present invention provide methods and systems useful for reprojection of images or videos from a viewpoint corresponding to a current head pose. The present invention is applicable to various applications in computer vision and image display systems.

[0012]

[0036] Reference is now made to the drawings, where like reference numerals refer to like parts throughout. Unless otherwise indicated, the drawings are schematic and are not necessarily drawn to scale.

[0013]

[0037] Referring now to FIG. 2A, in some embodiments, light striking the waveguide may need to be redirected so as to internally couple that light into the waveguide. An internal coupling optical element may be used to redirect and internally couple the light into its corresponding waveguide. Although referred to throughout this specification as an "internal coupling optical element," the internal coupling optical element need not be an optical element and may be a non-optical element. FIG. 2A shows a cross-sectional side view of an example of a set 200 of stacked waveguides each including an internal coupling optical element. Each waveguide may be configured to output light of one or more different wavelengths, or one or more different wavelength ranges. Light from a projector is input into the set 200 of stacked waveguides and is externally coupled to a user as more fully described below.

[0014]

[0038] The illustrated set of stacked waveguides 200 includes waveguides 202, 204, and 206. Each waveguide includes associated internally coupled optical elements (which may also be referred to as optical input areas on the waveguide), for example, an internally coupled optical element 203 disposed on the main surface (e.g., upper main surface) of waveguide 202, an internally coupled optical element 205 disposed on the main surface (e.g., upper main surface) of waveguide 204, and an internally coupled optical element 207 disposed on the main surface (e.g., upper main surface) of waveguide 206. In some embodiments, one or more of the internally coupled optical elements 203, 205, and 207 may be disposed on the bottom main surface of each waveguide 202, 204, and 206 (particularly when one or more internally coupled optical elements are reflective deflection optical elements). As shown in the figures, the internally coupled optical elements 203, 205, and 207 may be located on the upper main surface of their respective waveguides 202, 204, and 206 (or on the upper part of the subsequent lower waveguide), particularly if those internally coupled optical elements are transmissive deflection optical elements. In some embodiments, the internally coupled optical elements 203, 205, and 207 may be located within the body of their respective waveguides 202, 204, and 206. In some embodiments, as considered herein, the internally coupled optical elements 203, 205, and 207 are wavelength-selective, selectively redirecting light of one or more wavelengths while transmitting light of other wavelengths. Although the internally coupled optical elements 203, 205, and 207 are shown on one side or corner of their respective waveguides 202, 204, and 206, it will be recognized that in some embodiments they may be located in other areas of their respective waveguides 202, 204, and 206.

[0015]

[0039] As shown in the figure, the internally coupled optical elements 203, 205, and 207 may be offset laterally from each other. In some embodiments, each internally coupled optical element may be offset so that light does not pass through another internally coupled optical element. For example, each internally coupled optical element 203, 205, and 207 may be configured to receive light from different projectors and may be separated (e.g., laterally spaced) from the other internally coupled optical elements 203, 205, and 207 so as not to receive light from the other internally coupled optical elements 203, 205, and 207.

[0016]

[0040] Each waveguide also includes associated optical distribution elements, for example, an optical distribution element 210 disposed on the main surface (e.g., upper main surface) of waveguide 202, an optical distribution element 212 disposed on the main surface (e.g., upper main surface) of waveguide 204, and an optical distribution element 214 disposed on the main surface (e.g., upper main surface) of waveguide 206. In some other embodiments, the optical distribution elements 210, 212, and 214 may each be disposed on the bottom main surface of the associated waveguides 202, 204, and 206. In some other embodiments, the optical distribution elements 210, 212, and 214 may each be disposed on both the top and bottom main surfaces of the associated waveguides 202, 204, and 206. Alternatively, the optical distribution elements 210, 212, and 214 may be arranged on different surfaces of the upper and lower main surfaces of different associated waveguides 202, 204, and 206, respectively.

[0017]

[0041] Waveguides 202, 204, and 206 may be separated and isolated by layers of material, for example, gas, liquid, and / or solid. For example, as shown, layer 208 may separate waveguides 202 and 204, and layer 209 may separate waveguides 204 and 206. In some embodiments, layers 208 and 209 are formed from low refractive index materials (i.e., materials with a lower refractive index than the materials forming directly adjacent waveguides among waveguides 202, 204, and 206). Preferably, the refractive index of the materials forming layers 208 and 209 is 0.05 or more, or 0.10 or less, lower than the refractive index of the materials forming waveguides 202, 204, and 206. Advantageously, layers 208 and 209 with lower refractive indices can function as cladding layers that facilitate total internal reflection (TIR) ​​of light passing through waveguides 202, 204, and 206 (e.g., TIR between the upper and lower main surfaces of each waveguide). In some embodiments, layers 208 and 209 are formed from air. Although not shown, it will be recognized that the upper and lower parts of the illustrated set of waveguides 200 may include directly adjacent cladding layers.

[0018]

[0042] Preferably, for ease of manufacture and other considerations, the materials forming waveguides 202, 204, and 206 are similar or identical, and the materials forming layers 208, 209 are similar or identical. In some embodiments, the materials forming waveguides 202, 204, and 206 may differ between one or more waveguides, and / or the materials forming layers 208, 209 may differ while still maintaining the various refractive index relationships described above.

[0019]

[0043] Continuing to refer to Figure 2A, rays 218, 219, and 220 are incident on waveguide set 200. It will be recognized that rays 218, 219, and 220 may also be introduced into waveguides 202, 204, and 206 by one or more projectors (not shown).

[0020]

[0044] In some embodiments, the light rays 218, 219, and 220 have different characteristics, such as different wavelengths or different wavelength ranges, which may correspond to different colors. Each internally coupled optical element 203, 205, and 207 deflects the incident light so that it propagates through waveguides 202, 204, and 206 respectively by TIR. In some embodiments, each of the internally coupled optical elements 203, 205, and 207 selectively deflects one or more specific wavelengths of light, while allowing other wavelengths to pass through the underlying waveguide and associated internally coupled optical elements.

[0021]

[0045] For example, the internally coupled optical element 203 may be configured to deflect a ray 218 having a first wavelength or wavelength range, and to transmit rays 219 and 220 having different second and third wavelengths or wavelength ranges, respectively. The transmitted ray 219 collides with and is deflected by an internally coupled optical element 205 configured to deflect light of the second wavelength or wavelength range. The ray 220 is deflected by an internally coupled optical element 207 configured to selectively deflect light of the third wavelength or wavelength range.

[0022]

[0046] Continuing to refer to Figure 2A, the deflected rays 218, 219, and 220 are deflected so that they propagate through their corresponding waveguides 202, 204, and 206. That is, the internal coupling optical elements 203, 205, and 207 of each waveguide deflect the light into their corresponding waveguides 202, 204, and 206, thereby internally coupling the light within their corresponding waveguides. The rays 218, 219, and 220 are deflected by TIR at an angle that causes the light to propagate through their respective waveguides 202, 204, and 206. The rays 218, 219, and 220 propagate through their respective waveguides 202, 204, and 206 by TIR until they collide with the corresponding optical distribution elements 210, 212, and 214 of the waveguides, where they are externally coupled to provide the externally coupled ray 216.

[0023]

[0047] Referring now to Figure 2B, a perspective view of an example of the stacked waveguides in Figure 2A is shown. As described above, the internally coupled rays 218, 219, and 220 are deflected by the internally coupled optical elements 203, 205, and 207, respectively, and then propagate through waveguides 202, 204, and 206 by TIR, respectively. The rays 218, 219, and 220 then collide with the optical distribution elements 210, 212, and 214, respectively. The optical distribution elements 210, 212, and 214 deflect the rays 218, 219, and 220 so that they propagate toward the externally coupled optical elements 222, 224, and 226, respectively.

[0024]

[0048] In some embodiments, the light distribution elements 210, 212, and 214 are orthogonal pupil expanders (OPEs). In some embodiments, the OPEs deflect or distribute light to externally coupled optics 222, 224, and 226, and in some embodiments, they can also increase the beam or spot size of this light as it propagates to the externally coupled optics. In some embodiments, the light distribution elements 210, 212, and 214 may be omitted, and the internally coupled optics 203, 205, and 207 may be configured to deflect light directly to the externally coupled optics 222, 224, and 226. For example, referring to Figure 2A, the light distribution elements 210, 212, and 214 may be replaced by externally coupled optics 222, 224, and 226, respectively. In some embodiments, the externally coupled optics 222, 224, and 226 are exit pupils (EPs) or exit pupil expanders (EPEs) that direct light towards the user's eye. It will be recognized that the OPE may be configured to increase the dimensions of the eyebox on at least one axis, and the EPE may be configured to increase the eyebox on an axis intersecting the axis of the OPE, for example, an orthogonal axis. For example, each OPE may be configured to redirect a portion of the light that hits the OPE to the EPE in the same waveguide, while allowing the rest of the light to continue propagating through the waveguide. Upon hitting the OPE again, another portion of the remaining light is redirected to the EPE, and the rest of that portion continues propagating through the waveguide, and so on. Similarly, upon hitting the EPE, a portion of the colliding light is redirected out of the waveguide toward the user, and the rest of that light continues propagating through the waveguide until it hits the EPE again, at which point another portion of the colliding light is redirected out of the waveguide, and so on. As a result, a single beam of internally coupled light can be "duplicated" each time a portion of its light is redirected by the OPE or EPE, thereby forming a field of cloned light beams. In some embodiments, the OPE and / or EPE may be configured to resize the light beam. In some embodiments, the functions of the light distribution elements 210, 212, and 214 and the externally coupled optical elements 222, 224, and 226 are combined in a combined pupil expander, as discussed in relation to Figure 2E.

[0025]

[0049] Therefore, referring to Figures 2A and 2B, in some embodiments, the set of waveguides 200 includes waveguides 202, 204, 206 for each component color, internally coupled optics 203, 205, 207, optical distribution elements (e.g., OPE) 210, 212, 214, and externally coupled optics (e.g., EP) 222, 224, 226. Waveguides 202, 204, 206 may be stacked with air gaps / cladding layers between them. The internally coupled optics 203, 205, 207 redirect or deflect the incident light into their waveguides (by different internally coupled optics receiving light of different wavelengths). The light then propagates within each waveguide 202, 204, 206 at an angle that yields a TIR. In the illustrated example, ray 218 (e.g., blue light) is deflected by the first internal coupling optical element 203, then bounces along the waveguide, interacting with the optical distribution element (e.g., OPE) 210 and then the external coupling optical element (e.g., EP) 222 in the manner described above. Rays 219 and 220 (e.g., green light and red light, respectively) pass through waveguide 202, with ray 219 colliding with the internal coupling optical element 205 and being deflected. Ray 219 then bounces along waveguide 204 via TIR, proceeding to its optical distribution element (e.g., OPE) 212, and then to the external coupling optical element (e.g., EP) 224. Finally, ray 220 (e.g., red light) passes through waveguide 206 and collides with the optical internal coupling optical element 207 of waveguide 206. The internal optical coupling element 207 deflects the light ray 220 so that after the light ray propagates to the optical distribution element (e.g., OPE) 214 by TIR, it propagates to the external coupling optical element (e.g., EP) 226 by TIR. The external coupling optical element 226 then finally externally couples the light ray 220 to the viewer, who also receives externally coupled light from the other waveguides 202 and 204.

[0026]

[0050] Figure 2C is a top view of an example of the stacked waveguides shown in Figures 2A and 2B. As shown, waveguides 202, 204, and 206 may be vertically aligned with their associated optical distribution elements 210, 212, and 214 and associated external coupling optics 222, 224, and 226. However, as discussed herein, the internal coupling optics 203, 205, and 207 are not vertically aligned. Rather, the internal coupling optics are preferably non-overlapping (e.g., laterally spaced as seen in the top or plan view). As further discussed herein, this non-overlapping spatial arrangement facilitates one-to-one input of light from different resources to different waveguides, thereby enabling the unique coupling of a particular light source to a particular waveguide. In some embodiments, arrangements including non-overlapping, spatially separated internal coupling optics may be referred to as a pupil-shifting system, where the internal coupling optics in these arrangements can correspond to sub-pupils.

[0027]

[0051] Figure 3 is a simplified diagram of an eyepiece waveguide having a combined pupil expander according to one embodiment of the present invention. In the example shown in Figure 3, the eyepiece 310 utilizes a combined OPE / EPE region in a one-sided configuration. Referring to Figure 3, the eyepiece 310 includes a substrate 320 on which an internally coupled optical element 322 and a combined OPE / EPE region 324, also referred to as a combined pupil expander (CPE), are provided. The incident ray 330 is internally coupled via the internally coupled optical element 322 and externally coupled as an output ray 332 via the combined OPE / EPE region 324.

[0028]

[0052] The combined OPE / EPE region 324 includes grids corresponding to both OPE and EPE that spatially overlap in the x and y directions. In some embodiments, the grids corresponding to both OPE and EPE are located on the same side of the substrate 320 such that the OPE grid is superimposed on the EPE grid, or the EPE grid is superimposed on the OPE grid (or both). In other embodiments, the OPE grid is located on the opposite side of the substrate 320 from the EPE grid such that the grids spatially overlap in the x and y directions but are separated from each other in the z direction (i.e., in different planes). Thus, the combined OPE / EPE region 324 can be implemented in either a one-sided or two-sided configuration.

[0029]

[0053] Figure 4 shows an example of a wearable display system 430 in which various waveguides and associated systems disclosed herein may be integrated. Continuing with reference to Figure 4, the display system 430 includes a display 432 and various mechanical and electronic modules and systems to support the functionality of the display 432. The display 432 is wearable by a user 440 (also referred to as the viewer) of the display system and can be coupled to a frame 434 configured to position the display 432 in front of the user 440's eyes. In some embodiments, the display 432 may be considered eyewear. In some embodiments, a speaker 436 is coupled to the frame 434 and configured to be positioned adjacent to the user 440's ear canal (in some embodiments, another speaker, not shown, may be optionally positioned adjacent to the user's other ear canal to provide stereo / shapeable acoustic control). The display system 430 may also include one or more microphones or other devices for detecting sound. In some embodiments, the microphone is configured to allow the user to provide input or commands to the system 430 (e.g., selection of voice menu commands, natural language questions, etc.) and / or to enable voice communication with other people (e.g., with other users of a similar display system). The microphone may further be configured as a peripheral sensor for collecting voice data (e.g., sounds from the user and / or the environment). In some embodiments, the display system 430 may further include one or more outward-oriented environmental sensors configured to detect objects, stimuli, people, animals, places, or other aspects of the world around the user. For example, the environmental sensors may include one or more cameras, which may be positioned outward, for example, to capture images similar to at least a portion of the user 440's normal field of view.In some embodiments, the display system may also be separate from the frame 434 and may include peripheral sensors that can be attached to the user 440's body (e.g., the user 440's head, torso, limbs, etc.). In some embodiments, the peripheral sensors may be configured to acquire data characterizing the user 440's physiological state. For example, the sensors may be electrodes.

[0030]

[0054] The display 432 is operably coupled to a local data processing module by a communication link such as a wired lead wire or a wireless connection, and the local data processing module can be mounted in various configurations, such as being fixedly mounted to a frame 434, fixedly mounted to a helmet or hat worn by the user, embedded in headphones, or otherwise detachably mounted to the user 440 (e.g., in a backpack configuration, in a belt-mounted configuration). Similarly, sensors can be operably coupled to the local processor and data module by a communication link such as a wired lead wire or a wireless connection. The local processing and data module may include a hardware processor and digital memory such as non-volatile memory (e.g., flash memory or a hard disk drive), both of which can be used to assist in data processing, caching, and storage. Optionally, the local processor and data module may include one or more central processing units (CPUs), graphics processing units (GPUs), dedicated processing hardware, etc. The data may include a) data captured from sensors such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, a gyroscope, and / or other sensors disclosed herein (which may, for example, be operably coupled to frame 434 or otherwise attached to user 440), and / or b) data acquired and / or processed using a remote processing module 452 and / or a remote data repository 454 (including data relating to virtual content), possibly for passage to display 432 after such processing or retrieval. The local processing and data module may be operably coupled to the remote processing and data module 450 by a communication link 438, such as via a wired or wireless link, and the remote processing and data module may include the remote processing module 452, the remote data repository 454, and a battery 460.The remote processing module 452 and the remote data repository 454 can be coupled to the remote processing and data module 450 by communication links 456 and 458 so that these remote modules are operablely coupled to each other and available as resources to the remote processing and data module 450. In some embodiments, the remote processing and data module 450 may include one or more of the following: an image capture device, a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope. In some other embodiments, one or more of these sensors may be mounted on the frame 434 or may be a standalone structure communicating with the remote processing and data module 450 by a wired or wireless communication path.

[0031]

[0055] Continuing to refer to Figure 4, in some embodiments, the remote processing and data module 450 may include one or more processors configured to analyze and process data and / or image information, such as one or more central processing units (CPUs), graphics processing units (GPUs), and dedicated processing hardware. In some embodiments, the remote data repository 454 may include digital data storage facilities that may be available via the Internet or other networking configurations in a “cloud” resource configuration. In some embodiments, the remote data repository 454 may include one or more remote servers that provide information, such as information for generating augmented reality content, to the local processing and data module and / or the remote processing and data module 450. In some embodiments, all data is stored, all calculations are performed on the local processing and data module, and fully autonomous use from the remote module is possible. Optionally, an external system including a CPU, GPU, etc. (e.g., one or more processors, one or more computer systems) may perform at least part of the processing (e.g., generating image information, processing data) and provide and receive information from the illustrated module, for example, via a wireless or wired connection.

[0032]

[0056] Figure 5 shows a perspective view of a wearable device 500 according to one embodiment of the present invention. The wearable device 500 includes a frame 502 configured to support one or more projectors 504 at various positions along a surface facing the interior of the frame 502, as shown in the figure. In some embodiments, the projectors 504 can be mounted near the temples 506. Alternatively, or in addition, another projector can be placed at position 508. Such a projector may include, for example, one or more liquid crystal on silicon (LCoS) modules, micro-LED displays, or fiber scanning devices, or may operate in conjunction with them. In some embodiments, light from projectors 504 or projectors placed at position 508 can be directed into an eyepiece 510 for display to the user's eye. A projector placed at position 512 can be somewhat smaller due to the proximity this brings to the waveguide system. The closer the position, the less light can be lost when the waveguide system directs the light from the projector to the eyepiece 510. In some embodiments, the projector at position 512 can be used in conjunction with projector 504 or a projector positioned at position 508. Although not depicted, in some embodiments, the projector can also be positioned below the eyepiece 510. The wearable device 500 is depicted including sensors 514 and 516. Sensors 514 and 516 may take the form of forward-facing and lateral-facing optical sensors configured to characterize the real-world environment surrounding the wearable device 500.

[0033]

[0057] Embodiments of the present invention utilize an eye-tracking system to determine the user's gaze position and to utilize the gaze position for an image compression process. Referring to Figure 5, an eye-tracking camera 505 is positioned on frame 502 and can be used to track the user's gaze position using a wearable device 500. In other embodiments, the gaze position is determined using other eye-tracking systems, and the eye-tracking camera 505 shown in Figure 5 is merely illustrative. Among other functions, as will be described more fully herein, the image compression process, internal communications, and display used to compress and decompress virtual content for storage in memory can be modified according to the gaze position. For example, portions of an image or video stream corresponding to a gaze position can be compressed using a higher-quality compression process compared to other portions of the image or video stream located further away from the gaze position. Since these further portions of the image or video stream are in the user's peripheral vision, any impact on the user experience resulting from reduced compression quality may be smaller than the benefits achieved in terms of memory and processing efficiency and / or requirements. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0034]

[0058] In augmented reality (AR) systems where wearable devices overlay computer-generated images onto the existing world, image correction is performed to provide consistent fixation characteristics to the virtual image. Therefore, when the virtual image is placed in a real-world location, the image preferably does not move or jitter relative to its real-world placement. This characteristic can be referred to as pixel sticking.

[0035]

[0059] During the use of an AR system, image correction is performed because the position of the headset may change from the time the image is generated until the final image is displayed. Therefore, the computing system that generates the original display image (e.g., a GPU) can also predict the future position of the headset device to reduce or minimize errors caused by headset movement. Some AR systems also perform a final additional correction based on the actual headset position before displaying the image on the headset.

[0036]

[0060] In some implementations, the processor that initially generates the content is located close to the headset. However, cloud-based implementations can perform cloud-based rendering, also known as remote rendering, instead of using a local computing system. In these cloud-based implementations, the headset prediction process can be severely affected by transmission latency that occurs during data communication.

[0037]

[0061] Remote rendering systems can have extremely high latency, sometimes requiring a complete reprojection. This reprojection process, known as depth-based reprojection, typically utilizes a significant amount of computational power on traditional GPUs and results in considerable power consumption.

[0038]

[0062] In wireless systems, MPEG encoder and decoder systems can be used to reduce radio frequency (RF) bandwidth. In this system, a remote rendering system can MPEG encode the incoming video stream, and then a wearable unit can decode the video stream. Additionally, this system typically implements late 3DOF or 6DOF display correction (WARP) to compensate for the difference between the current head pose when the image in the video stream is displayed and the predicted head pose corresponding to the rendered image. Thus, the remote rendering system predicts the future head pose position and then renders the image from that viewpoint.

[0039]

[0063] Therefore, to compensate for latency in wireless systems, 6DOF display correction is implemented according to embodiments of the present invention. This correction uses a depth map, also called a depth-based map, along with a color map, to reproject the image from a different viewpoint, i.e., the current head pose.

[0040]

[0064] Figure 6 is a simplified diagram illustrating the reprojection of virtual content according to one embodiment of the present invention. This figure shows a simulation of a change in viewpoint over time by warping frames. This process can be referred to as time warping. Referring to Figure 6, the compositor frame 610 contains 3D virtual content 605 generated from the viewpoint of the compositor viewpoint 612. The compositor frame 610 can represent the right image of a right / left image pair. Alternatively, the compositor frame 610 can represent the left image of a right / left image pair.

[0041]

[0065] Therefore, referring to Figure 6, one of the multiple views shown by the uncompressed virtual content is represented by the combiner frame 610. The virtual content within the combiner frame 610 is generated from the combiner viewpoint 612, but the viewpoint from which the virtual content should be displayed to the user is the client viewpoint 622. Thus, the 3D virtual content 605 is warped or reprojected from the client viewpoint 622 in the form of the client frame 620. Thus, referring to Figure 6, one of the multiple views shown by the uncompressed content is represented by the client frame. Using an embodiment of the present invention, for example, the image reprojection system 700 shown in Figure 7, the uncompressed virtual content (e.g., the combiner frame 610) is generated for a predicted head pose corresponding to the client viewpoint 622, but is reprojected as uncompressed content (e.g., the client frame 620) corresponding to the current head pose using an MV-HEVC encoder 720 and an MV-HEVC decoder 740 that take the current head pose or adjustment 742 as input.

[0042]

[0066] Figure 7 is a simplified schematic diagram showing an image reprojection system according to one embodiment of the present invention. Referring to Figure 7, the image reprojection system 700 is capable of receiving uncompressed content 712 corresponding to a given pose, which is referred to herein as a predicted head pose. The uncompressed content may be, for example, virtual content generated using a GPU. Each of the generated uncompressed images may be represented by a depth map and a corresponding color map. As an example, multiple views may be generated, for example, left and right images or left and right video streams for display to a user using an AR headset. Thus, in some implementations, these multiple video streams (also referred to as multiple views, since each video stream may correspond to a different pose) may be generated using a GPU. In some embodiments, a single depth map and a corresponding color map are generated corresponding to the predicted head pose. Then, using this single depth map and corresponding color map, left and right images are generated from either the predicted head pose or the adjusted head pose, as will be discussed more fully herein. Thus, embodiments of the present invention are not limited to uncompressed content 712 containing multiple images or video streams, since given a single image or video stream and a corresponding pose, additional images or video streams can be appropriately generated. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0043]

[0067] Additionally, uncompressed content 712 can be generated by multiple cameras capturing multiple views of a scene. Combinations can be provided that capture images from a given pose and create virtual content based on those images to provide multiple images corresponding to multiple views. Embodiments of the present invention are not limited to images, and video streams are included within the scope of the invention.

[0044]

[0068] Therefore, the uncompressed content 712 includes virtual content corresponding to multiple poses, images or videos captured by multiple cameras with different poses, and combinations thereof. Although the uncompressed content 712 is considered herein and shown in the context of a set of left and right images or a video stream, it will be recognized that a single video stream corresponding to a given pose can be generated and is included within the scope of the embodiments described herein.

[0045]

[0069] In addition to uncompressed content 712, a predicted head pose or predicted content pose 710 is provided. The predicted head pose or predicted content pose 710, as well as the uncompressed content 712, can be generated remotely from the AR headset, for example, by a cloud service, or by a combination of these sources, in another processing device connected to the beltpack or AR headset. Given the predicted head pose or predicted content pose 710, the MV-HEVC encoder 720 performs view coding, depth coding, and color coding to generate compressed content 730, which can be made into a compressed video stream. The MV-HEVC encoder 720 can be implemented on the user's AR headset or associated equipment, or remotely using, for example, cloud-based computing resources.

[0046]

[0070] As shown in Figure 7, the predicted head pose or predicted content pose 710 can be implemented using the predicted content pose. As an example, as shown in Figure 6, the client frame 620 can be generated from the client viewpoint 622, and this dataset can be provided instead of the predicted head pose. Given the client frame 620 and the client viewpoint 622, a reprojection can be performed to reproject the synthesizer frame 610 from the synthesizer viewpoint 612. Thus, as an extension of existing MV-HEVC encoding processes, embodiments of the present invention encode virtual content in the context of the predicted content pose or head pose (e.g., client viewpoint) as well as virtual content (i.e., left and right views). The decoding stage can then correct the virtual content to correspond to the actual head pose at display. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0047]

[0071] As will be explained more fully in relation to Figure 22, a spherical depth map can be generated instead of the predicted head pose or predicted content pose 710. In this case, virtual content can be generated from a given pose. The virtual content can be mapped onto the spherical depth map, and the spherical depth map and the given pose can be provided in the same way as the predicted head pose. Given the reprojection pose, the spherical depth map, and the given pose, reprojection can be performed in accordance with the display of the virtual content from the viewpoint of the reprojection pose.

[0048]

[0072] Referring again to Figure 7, the compressed content 730 (e.g., compressed video) is provided to the MV-HEVC decoder 740 as input. In addition, the MV-HEVC decoder 740 receives not only the compressed content 730 as input but also the current head attitude or adjustment 742 from the headset or other preferred source. In some embodiments, an absolute current head attitude, also referred to as the actual current head attitude, is utilized in the decoding process. In other embodiments, the adjustment is calculated as the difference (i.e., Δ) between the predicted head attitude and the current head attitude and is utilized in the decoding process. Thus, either an adjustment to the actual head attitude or a predicted head attitude can be utilized by the embodiments described herein.

[0049]

[0073] Using a depth map, color map, and current head pose or adjustment 742, the MV-HEVC decoder 740 implements an MV-HEVC decoding process that generates uncompressed head pose corrected content 750, which, like uncompressed content 712, can contain multiple image or video streams and can therefore be referred to as head pose corrected multiple views. Thus, the uncompressed head pose corrected content 750, for example, GPU-generated virtual content corresponding to multiple views, is reprojected based on the current head pose or adjustment to generate uncompressed head pose corrected content 750. In some implementations, the multiple views generated by the MV-HEVC decoder 740 are left and right views suitable for display in the left and right eyepieces of an AR headset. By decoding the compressed content 730 so that the viewpoint corresponds to the current head pose, the uncompressed content 712 (e.g., left and right streams of virtual content) is reprojected to match the user's head pose at the time the virtual content is displayed (i.e., updated with respect to predicted head pose or predicted content pose) and can be characterized by a desired pixel stick. As a result, embodiments of the present invention improve the user experience by compensating for the latency within the system between the time when virtual content is generated and the time when the virtual content is displayed to the user, thereby increasing the pixel stick.

[0050]

[0074] In some embodiments, encoding by the MV-HEVC encoder 720 is performed at a first frame rate, for example, 60 Hz. However, decoding by the MV-HEVC decoder 740 can be performed at a second frame rate higher or lower than the first frame rate. For example, decoding by the MV-HEVC decoder 740 can be performed at frame rates such as 120 Hz, 240 Hz, or 360 Hz. Furthermore, using the current head pose or adjustment 742, an uncompressed image is generated from the viewpoint of the current head pose. For example, if uncompressed content 712 is generated for the initial head pose, e.g., a head pose centered on the origin of a Cartesian coordinate frame and oriented along an axis, and the current head pose or adjustment 742 indicates that the head pose has changed to the current head pose, which is still centered on the origin but rotated by a predetermined angle with respect to the axis, then uncompressed head pose corrected content 750 (e.g., uncompressed head pose corrected multiple views) is generated from the viewpoint of the current head pose, thereby performing the final stage of 6DOF image reprojection correction, which ensures that the virtual content is accurately positioned in the world view from the user's viewpoint. Thus, it is possible to generate a stream of uncompressed images from multiple views that match the current head pose.

[0051]

[0075] Figure 8 is a simplified flowchart illustrating a method for encoding and decoding an image based on head pose according to one embodiment of the present invention. Method 800 comprises an encoding process shown by steps 810-826 and a decoding process shown by steps 830-840. Method includes receiving images from multiple cameras (i.e., N cameras) or one or more depth map / color map combinations (810). Depth map / color map combinations can be generated using a GPU or other suitable processor. Method also includes obtaining a predicted head pose (820). Depth map / color map combinations can correspond to the predicted head pose. Method further includes performing depth estimation to create a depth / color map of a desired projection view corresponding to the predicted head pose (822). The depth / color maps are compressed using a desired multiview encoding standard (824) to produce a compressed video (826).

[0052]

[0076] During the decoding process, the method includes receiving compressed video (830) and decompressing the depth / color map using a desired multiview decoding standard (832). The method may receive the current head pose (836) and calculate the difference between the predicted head pose and the current head pose to provide reprojection correction (834). In some embodiments, the current head pose is used to perform the reprojection and the difference is not used. The method includes recreating a combination of depth / color maps or multiple camera views (e.g., left and right views) from the viewpoint of the current head pose (838). The method may also include displaying the left and right images or views on left and right displays, respectively (840).

[0053]

[0077] During use of an AR headset, user movement may cause the current head posture not to always match the predicted head posture. Therefore, embodiments of the present invention can reproject virtual content from the predicted head posture to the current head posture to improve the user experience. By using the MV-HEVC process, significant reductions in processor size, power consumption, and capabilities can be achieved.

[0054]

[0078] Some embodiments of the present invention implement a higher head pose correction frame rate, as shown in Figure 8. In these embodiments implementing a higher head pose correction frame rate, the current head pose can be received (836) after the left and right images are displayed on the left and right displays respectively (840), and view reprojection correction is performed (834) using the calculated difference between the current or predicted head pose and the current head pose (i.e., adjustment), and a combination of depth maps / color maps or multiple camera views (e.g., left and right views) from the viewpoint of the current head pose is recreated (838), and the left and right images or views are displayed on the left and right displays respectively (840). This loop (i.e., steps 834, 838, and 840) can be performed at a higher frame rate than the encoding and decoding process. Therefore, for example, if the virtual content generation, encoding, and decoding processes are operating at 60Hz, but the current head pose (836) is generated at a rate of 120Hz, the loop (i.e., steps 834, 838, and 840) can operate at 120Hz, updating the virtual content on the display twice based on the current head pose for each frame of the generated virtual content. For example, if the current head pose is available at a rate of 180Hz, the virtual content can be updated and displayed by performing head pose correction twice for each frame of the generated virtual content. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0055]

[0079] It should be noted that the specific steps shown in Figure 8 provide a specific method for encoding and decoding an image based on head pose, according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figure 8 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0056]

[0080] As described above, in some embodiments, the encoder is provided with a predicted head pose, and therefore the encoder has the approximate position of the remote headset and renders a depth and color map of its predicted head pose. The decoder then uses the current head pose to appropriately reproject the image(s) at the desired viewpoint. This makes it possible to present the virtual content in the correct position even if the headset may have drifted from the desired position.

[0057]

[0081] In another embodiment, the encoder generates a spherical depth map with a world-lock reference, as considered in relation to Figure 22. The headset can then use the given reference to reproject the correct viewpoint from any desired position around the sphere.

[0058]

[0082] Figure 9 is a simplified schematic diagram illustrating the operation of an image reprojection system according to one embodiment of the present invention. Referring to Figure 9, the image reprojection system 900 is capable of receiving compressed content corresponding to a given orientation using one or more communication interfaces 910. As discussed above in relation to Figure 7, the compressed content received by the image reprojection system 900 can be generated using an MV-HEVC encoder that can be located remotely from the image reprojection system 900, for example, using cloud-based computing resources. In the illustrated embodiment, the compressed content is received using WiFi, USB, DisplayPort (DP), or other communication protocols. In this embodiment, the compressed content is MPEG video.

[0059]

[0083] Compressed content can be, for example, virtual content generated using a GPU. Each of the generated compressed images can be represented by a depth map and a corresponding color map. As an example, multiple views can be generated, for example, left and right images or left and right video streams for display to a user using an AR headset. Thus, in some implementations, these multiple video streams (also referred to as multiple views, since each video stream can correspond to a different pose) can be generated using a GPU. In some embodiments, a single depth map and a corresponding color map are generated corresponding to a predicted head pose. Then, using this single depth map and corresponding color map, left and right images are generated from either the predicted or adjusted head pose, as will be discussed more fully herein. Thus, embodiments of the present invention are not limited to compressed content containing multiple images or video streams, since given a single image or video stream and a corresponding pose, additional images or video streams can be appropriately generated. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0060]

[0084] Additionally, compressed content can be generated by multiple cameras capturing multiple views of a scene. Combinations can be provided that capture images from a given pose and create virtual content based on those images to provide multiple images corresponding to multiple views. Embodiments of the present invention are not limited to images; video streams are included within the scope of the invention.

[0061]

[0085] Therefore, compressed content includes virtual content corresponding to multiple poses, images or videos captured by multiple cameras with different poses, and combinations thereof. Although compressed content is considered herein and shown in the context of a set of left and right images or a video stream, it will be recognized that a single video stream corresponding to a given pose can be generated and is included within the scope of the embodiments described herein.

[0062]

[0086] Additionally, the image reprojection system 900 can receive predicted content pose or predicted head pose using one or more motion detection devices 920. In the illustrated embodiment, the predicted content pose or predicted head pose is received using an inertial motion unit (IMU), a head pose tracking system, a head tracking system, or an eye-tracking system. The predicted head pose or predicted content pose, as well as the compressed content, can be generated remotely from the AR headset, for example, by a cloud service or a combination of these sources, in another processing device connected to the beltpack or AR headset. The predicted content pose or predicted head pose is received by a central processing unit (CPU) / neural processing unit (NPU) controller 930. The head pose is provided internally to the MV-HEVC decoder 940, and the gaze is provided internally to the foveal and spatial compression processor 950. This gaze information can then be used to perform fovealized image compression based on the gaze. This fovealized image compression process is described in more detail below in relation to Figures 11-14. In embodiments utilizing foveal and / or other spatial compression algorithms, foveal decompression and / or decompression can be performed on an external display 970 or using other suitable hardware and / or software.

[0063]

[0087] In some embodiments, line-of-sight based foveal projection is not implemented, and in these embodiments, the left-view and right-view views are passed through to produce a stereo output 960 that is displayed using an external display 970.

[0064]

[0088] Referring again to Figure 9, the compressed content (e.g., compressed video) is provided to the MV-HEVC decoder 940 as input. In addition, the MV-HEVC decoder 940 receives not only the compressed content as input from the CPU / NPU controller 930, but also the current head pose or adjustments to the head pose. In some embodiments, the absolute current head pose, also referred to as the actual current head pose, is used in the decoding process. In other embodiments, the adjustment is calculated as the difference (i.e., Δ) between the predicted head pose and the current head pose and is used in the decoding process. Thus, either the actual head pose or the adjustments to the predicted head pose can be used by the embodiments described herein.

[0065]

[0089] Given compressed content and a predicted head pose or predicted content pose, the MV-HEVC decoder 940 decodes the left viewpoint view stored in memory 942 and the right viewpoint view stored in memory 944. The MV-HEVC decoder 940 implements an MV-HEVC decoding process that generates uncompressed head pose corrected content, which, like the original compressed content, may contain multiple image or video streams, which may therefore be referred to as multiple head pose corrected views.

[0066]

[0090] Therefore, the uncompressed head pose correction content generated by the MV-HEVC decoder 940 is reprojected based on the current head pose or adjustments to the head pose to produce uncompressed head pose correction content. In some implementations, the multiple views generated by the MV-HEVC decoder 940 are left and right views suitable for display in the left and right eyepieces of the AR headset. By decoding the virtual content so that the viewpoint corresponds to the current head pose, the uncompressed content (e.g., left and right streams of virtual content) is reprojected to match the user's head pose at the time the virtual content is displayed (i.e., updated with respect to the predicted head pose or predicted content pose) and can be characterized by a desired pixel stick. As a result, embodiments of the present invention compensate for latency in the system between the time the virtual content is generated and the time the virtual content is displayed to the user, thereby improving the user experience by increasing the pixel stick.

[0067]

[0091] The foveal and spatial compression processor 950 receives a left-view perspective from memory 942 and a right-view perspective from memory 944, and generates a stereo output 960 which is displayed using an external display 970. In the illustrated embodiment, the stereo output 960 is a Mobile Industrial Processor Interface (MIPI) output, but other data formats are also within the scope of the present invention.

[0068]

[0092] Therefore, using the left and right viewpoint views generated using head posture, the foveal and spatial compression processor 950 can compress virtual content based on the user's gaze.

[0069]

[0093] Referring again to Figure 9 and the foveal and spatial compression processor 950, the line-of-sight based foveal process is shown in relation to Figures 11-14.

[0070]

[0094] Figure 10 is a diagram showing an image compressed using a single quality setting. In this case, all pixels in the image are compressed using a conventional process that utilizes a single quality setting for each pixel. While this process achieves uniform image compression across the image, the inventors determined that processing and memory requirements can be reduced if portions of the image farther from the user's viewing position are compressed at a reduced quality compared to the portion of the image corresponding to the user's viewing position, while still achieving the desired user experience.

[0071]

[0095] Figure 11 is a diagram showing a fovealed image having three fovealed regions according to one embodiment of the present invention. The image in Figure 11 is divided into multiple regions based on the gaze position. In this case, the user is fixated on the center of the image, and as a result, the gaze position is located at the center of the image. As discussed herein, the gaze position can be determined using an eye-tracking system such as those discussed in relation to Figures 5 and 22. Thus, the image can be divided into a central region corresponding to the gaze position and peripheral regions further away from the gaze position. In some embodiments, a foveal map is created based on the gaze position, with portions of the image closer to the gaze position mapped to a high-quality setting and portions of the image further away from the gaze position mapped to a low-quality setting. In Figure 11, the foveal map takes the form of two peripheral regions with lower quality settings and a central region with a higher (e.g., 100%) quality setting.

[0072]

[0096] In the image shown in Figure 11, region 1110, corresponding to the left quarter of the image (i.e., left 1 / 4), is compressed using a first quality setting. Additionally, region 1130, corresponding to the right quarter of the image (i.e., right 1 / 4), is compressed using a first quality setting. However, region 1120, corresponding to the middle half of the image (i.e., central 2 / 4), is compressed using a second quality setting, which is higher than the first quality setting. This division of the image into parts can be referred to as a three-region division: the left quarter (e.g., fovealed with a 70% quality setting), the middle half (e.g., non-fovealed with a 100% quality setting), and the right quarter (e.g., fovealed with a 70% quality setting).

[0073]

[0097] Figure 11 shows the division into three regions using a foveal map containing these three regions, but the present invention is not limited to this implementation, and images can be divided in other ways. By dividing an image into multiple regions, the quality setting of individual blocks or tiles (e.g., 8x8 pixel blocks for JPEG compression) contained in each region can be set to a predetermined quality setting for each block. Thus, in Figure 11, the same quality setting is assigned to all blocks within each region, i.e., the blocks in region 1110 are assigned a first quality setting (e.g., 70%), the blocks in region 1120 are assigned a second quality setting (e.g., 100%), and the blocks in region 1130 are assigned a first quality setting (e.g., 70%), but this is not mandatory, and different quality settings can be assigned to individual blocks within a region. Thus, the foveal map can be more complex than the three-region division shown in Figure 11. In some embodiments, blocks in the peripheral regions are assigned a quality setting that depends on the distance of the block from the line of sight, while blocks in the central region have a uniform quality setting. In other embodiments, the foveal map can be defined such that blocks in the peripheral region are assigned a uniform quality setting, while blocks in the central region are assigned a quality setting that depends on the distance of the block from the line of sight. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0074]

[0098] In the three-region fovealization image shown in Figure 11, an overall reduction of approximately 67% in image / memory size was achieved while maintaining 100% quality in region 1120, i.e., the non-fovealized section. As discussed above, the non-fovealized region (i.e., compressed using an uncompressed or lossless compression algorithm) can be any region identified in the foveal map. Consequently, the three-region division shown in Figure 11 is merely illustrative.

[0075]

[0099] It should be noted that if the gaze position is, for example, on the right side of the image, the foveal map can compress the right side using a higher quality setting and the left side of the image using a lower quality setting. Therefore, in this example, if the gaze position is within region 1130, regions 1110 and 1120 are compressed using a first quality setting, and region 1130 is compressed using a second quality setting higher than the first. In some embodiments, for example, if the gaze position is within region 1130, region 1130 can be compressed using a higher quality setting, e.g., lossless compression; region 1120 can be compressed using an intermediate quality setting lower than the higher quality setting; and region 1110 can be compressed using a minimum quality setting lower than the intermediate quality setting. As a result, the fovea of ​​the image is a function of the gaze position, compressing or encoding the region containing the gaze position with a higher quality setting than one or more regions further away from the gaze position. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0076]

[0100] Furthermore, although Figure 11 shows a set of vertical regions, this is not essential to embodiments of the present invention, and the definition of regions can be carried out in other ways, including horizontally oriented regions, regions defined based on the distance to the line of sight, for example, a set of regions defined radially.

[0077]

[0101] Figure 12 is a foveated 3D generated image having three foveated regions according to yet another embodiment of the present invention. In Figure 12, the regions are defined similarly to those shown in Figure 11. However, in 3D generated images, since the majority of the image is black, much higher compression can be achieved. Using the method described herein, an 87% compression was achieved while maintaining 100% quality at the center of the image corresponding to the gaze position. In this example, region 1220 was compressed using a 100% quality setting (non-foveated at 100% quality setting), while regions 1210 and 1230 were compressed with a lower quality setting (foveated at 20% quality setting). Since in many examples of virtual content the image content is highest near the gaze position and the surrounding regions are dark or black, embodiments of the present invention are particularly well suited for use in virtual reality and augmented reality implementations.

[0078]

[0102] In some examples, all areas of an image can be compressed using a lower quality setting, while non-foveated areas can be compressed using a higher quality setting. Using the example in Figure 11, areas 1110, 1120, and 1130 can be compressed, respectively, using the lower quality setting for the foveated areas. Area 1120 can also be compressed using the high quality setting. When decoding a compressed image (for example, for reconstruction for display to a user), it may be desirable to decode sections of the image in parallel. Therefore, two decoders can be used to decode a compressed image. During image reconstruction, the decoded area 1120 using the high quality setting can be superimposed on the decoded areas 1110, 1120, and 1130 (i.e., the entire image) using the lower quality setting. The encoding may be JPEG (for example, using the quality settings described above), or it may be a technique including DSC or VDC-X (for example, using compression ratios), which are discussed more fully herein.

[0079]

[0103] Figure 13 is a diagram showing an image that can be used with multiple foveal maps according to one embodiment of the present invention. Figure 13 shows an image that includes a person 1306 located in section 1310, a tree 1302 located in sections 1320, 1322, 1330, and 1332, and a house 1304 located in sections 1324, 1326, 1338, and 1340. Different foveal maps can be created based on this image depending on the line of sight.

[0080]

[0104] If the user's gaze position is located in one of sections 1320, 1322, 1330, or 1332, i.e., if the user is looking at tree 1302, a foveal map can be used, and blocks in sections 1320, 1322, 1330, and 1332 are compressed using a 100% quality setting (non-fovealed at 100% quality), while blocks in the remaining sections (i.e., sections 1310, 1312, 1314, 1316, 1324, 1326, 1328, 1334, 1336, 1338, 1340, and 1342) are compressed using a lower quality setting (fovealed at 70% quality). Thus, image compression can be implemented using a foveal map that maintains quality within the region of the image corresponding to the gaze position, while peripheral parts of the image can be compressed using a lower quality setting to save system resources, including memory and processing.

[0081]

[0105] Alternatively, if the user's gaze position is in one of sections 1324, 1326, 1338, or 1340, i.e., if the user is looking at house 1304, the foveal map can be utilized, and the blocks in sections 1324, 1326, 1338, and 1340 are compressed using a 100% quality setting (non-fovealed at 100% quality), while the blocks in the remaining sections (i.e., sections 1310, 1312, 1314, 1316, 1320, 1322, 1328, 1330, 1332, 1334, and 1336, as well as 1342) are compressed using a lower quality setting (fovealed at 70% quality).

[0082]

[0106] Finally, when the user's gaze position is in section 1310, i.e., when the user is looking at person 1306, the foveal map can be utilized, and the blocks in section 1310 are compressed using a 100% quality setting (non-fovealed at 100% quality), while the blocks in the remaining sections (i.e., sections 1312, 1314, 1316, 1320, 1322, 1324, 1326, 1328, 1330, 1332, 1334, and 1336, 1338, 1340, and 1342) are compressed using a lower quality setting (fovealed at 70% quality). In some embodiments, the quality setting used for the remaining sections varies, for example, as a function of the distance from the gaze position. In these embodiments, the blocks in sections 1312, 1314, and 1316 can be compressed using a 90% quality setting, the blocks in sections 1320, 1322, 1324, 1326, and 1328 can be compressed using an 80% quality setting, and the blocks in sections 1330, 1332, 1334, and 1336, 1338, 1340, and 1342 can be compressed using a 70% quality setting. In some examples, instead of encoding with JPEG (e.g., using the quality settings described above), sections 1310-1342 may be compressed using techniques including DSC or VDC-X (e.g., using a compression ratio). For example, based on the gaze position, a non-tile-based compression technique such as DSC can be used to compress sections closer to the gaze position at a lower compression ratio, while sections further away from the gaze position can be compressed at a higher compression ratio.

[0083]

[0107] Figure 14 is a simplified flowchart illustrating a method for compressing an image according to one embodiment of the present invention. Method 1400 includes receiving an image (1410), determining the user's gaze position (1412), and generating a foveal map based on the gaze position (1414).

[0084]

[0108] The image may be an image contained within a video stream. Determining the user's gaze position can be achieved by utilizing an eye-tracking system that provides the gaze position as a function of time. The foveal map defines the quality at which blocks are compressed, which varies as a function of their position in the image; blocks in regions closer to the gaze position are compressed using a higher quality setting, and blocks in regions further away from the gaze position are compressed using a lower quality setting. In the example shown in Figure 11, the foveal map includes three regions, but the present invention is not limited to this particular implementation, and two or more regions can be defined. Furthermore, blocks within a given region can be compressed using a uniform quality setting, or they can be compressed with different quality settings depending on the particular implementation. In some embodiments, the foveal map includes a first region of the image and a second region of the image.

[0085]

[0109] The method also includes compressing a first region of an image using a first quality setting and compressing a second region of the image using a second quality setting (1416). In some embodiments, the first quality setting is an uncompressed quality setting or a lossless compression quality setting. Thus, blocks within the first region are compressed with a higher quality than other parts of the image. The second quality setting is a lower quality setting, e.g., a 70% quality setting, which reduces the data corresponding to the compressed image within these regions. As discussed above, since the user's line of sight places these regions in the user's peripheral vision, any loss of quality is offset by savings in memory and processor usage. The data compression processes for the first and second regions can be performed sequentially or in parallel, depending on the specific application.

[0086]

[0110] A compressed image or video, which may be referred to as a foveal image or video, can be transmitted to a display system along with the foveal map (1418), or stored in memory along with the foveal map (1419).

[0087]

[0111] In embodiments in which a compressed image or video is stored in memory together with a foveal map, method 1400 includes retrieving the fovealized image and foveal map from memory (1420), decompressing a first region of the image using a first quality setting, and decompressing a second region of the image using a second quality setting (1440). In embodiments in which a compressed image or video is transmitted to a display system together with a foveal map, method 1400 includes receiving the fovealized image and foveal map (1420), decompressing a first region of the image using a first quality setting, and decompressing a second region of the image using a second quality setting (1440). The decompression processes for the first and second regions can be performed sequentially or in parallel, depending on the specific application. The two regions can be merged to form a final image suitable for display (1442). The final image is then displayed on a display device (1444).

[0088]

[0112] It should be noted that the specific steps shown in Figure 14 provide a particular method for compressing an image according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figure 14 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0089]

[0113] Although the embodiments described above utilize a tile-based (also known as block-based) JPEG compression algorithm, embodiments of the present invention are not limited to this particular compression standard, and other compression standards can be used in conjunction with various embodiments of the present invention. As an example, Figures 15 to 21 describe a technique that uses run-length coding with DSC and VDC-X to compress video data.

[0090]

[0114] Figure 15 shows the compression levels obtained as a function of time, expressed by consecutive frames versus frequency, for both a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. In Figure 15, each frame was compressed using either a mask-based compression method or DSC, according to alternating algorithms that implement either a mask-based compression method or fully fixed-frame compression, such as DSC.

[0091]

[0115] As shown in Figure 15, each frame is analyzed to determine the number of lines with pixels that have a brightness level below a threshold. If the mask-based compression method results in a compression level greater than the compression threshold (e.g., 37%), the frame is compressed using the mask-based compression method. In Figure 15, this results in the first approximately 3800 frames being compressed using the mask-based compression method.

[0092]

[0116] When a mask-based compression method produces compressed frames with a compression level of less than 37%, such as frames with little black content, the DSC method is used. This results in these frames having a compression value of 37%. Referring to Figure 15, frames represented by a compression value of less than 37% are compressed using DSC, effectively baseline the minimum compression at 37%. Therefore, the frames in sets A and B have a compression value of 37%, rather than a lower value achieved using the mask-based compression method.

[0093]

[0117] Figure 16 shows a histogram of frame count versus compression for a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. As shown in Figure 16, the number of frames with less than approximately 37% compression is reduced to zero because the mask-based compression method was used for frames that could be compressed to a compression level greater than 37%, or the frame-based compression method (e.g., DSC) was used for the remaining frames that could not be compressed to a compression level greater than 37% using the mask-based compression method. Thus, the mask-based compression method operating alone produced some frames with a compression level of less than 37%, but the alternating method provided by embodiments of the present invention limits the minimum compression level to approximately 37%, as shown in Figure 16. For frames with significant black pixel content, the mask-based compression method provides a high level of compression, but for frames with limited black pixel content, the frame-based compression method establishes a lower limit of compression level, for example 37% in this illustrated embodiment. As will be apparent to those skilled in the art, the minimum compression level does not have to be 37%, which is merely illustrative, and other minimum compression levels can be utilized depending on the specific application. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0094]

[0118] Information about the compression method used for each frame can be provided to the endpoint, such as a decoder or display, so that the endpoint can utilize the appropriate decompression method when reconstructing each frame.

[0095]

[0119] Figure 17 is a simplified flowchart illustrating a method for compressing an image frame using an alternating compression algorithm according to one embodiment of the present invention. Method 1700 includes receiving a frame of video data (1710). The method also includes determining the number of lines in the frame that have pixel groups characterized by luminance levels below a threshold (1712).

[0096]

[0120] If the number of lines is greater than or equal to the compression threshold (1714), the frame is compressed using a mask-based compression method (1720). If the number of lines is less than the compression threshold, the frame is compressed using a frame-based compression method (1722). If additional frames exist (1730), the method operates on the next frame of video data by receiving the frame of video data (1710). Otherwise, the method terminates (1740). Thus, embodiments of the present invention alternate between compressing each frame using the respective compression methods, depending on the level of compression that can be achieved by each compression method.

[0097]

[0121] It should be noted that the specific steps shown in Figure 17 provide a particular method for compressing image frames using an alternating compression algorithm according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figure 17 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0098]

[0122] According to some embodiments of the present invention, for each frame, there is an embedded image line control or alternative control mechanism that provides the endpoint display with information on which system should be used to decode the incoming MIPI frame. In addition, a virtual MIPI channel can be used to indicate the compression ratio used by the endpoint display.

[0099]

[0123] Some embodiments of the present invention modify the compression quality based on target tracking, thereby reducing the quality in the foveal region to give a higher compression ratio. This is done for the MIPI interface, thereby reducing the amount of data transmitted to the LCOS / μLED display via MIPI. Thus, the embodiments also result in power saving.

[0100]

[0124] Embodiments of the present invention reduce the amount of stream-based data transmitted via MIPI compression. Furthermore, embodiments modify the compression quality based on target tracking, thereby providing a higher compression ratio to the foveal region where quality is reduced. Moreover, embodiments enable a higher compression ratio for stream-based compression techniques while maintaining quality in the area observed by the user. As a result, embodiments enable a much higher compression ratio while maintaining quality.

[0101]

[0125] For stream-based compression standards such as DSC and VESA display compression (VDC-X), low-latency implementations are utilized. This low-latency response is used so that any previous spatial warp adjustments performed remain applicable.

[0102]

[0126] Figure 18 is a simplified image showing an image frame divided into a high-quality region and a low-quality region according to one embodiment of the present invention. The image 1800 shown in Figure 18 includes a high-quality region 1810 and a low-quality region 1820. As will be discussed in more detail below, the high-quality region 1810 is compressed and decompressed using a first quality setting or compression level, and the low-quality region 1820 or the entire image is compressed and decompressed using a second quality setting or compression level, providing memory savings and other advantages. As an example, a single decoder can be utilized by not compressing the high-quality region 1810 and compressing the low-quality region using a single decoder. Significant savings can be achieved if the high-quality region 1810 is small compared to the entire image. An additional explanation regarding resizing the high-quality region is provided in U.S. Provisional Patent Application No. 63 / 543,876, filed October 12, 2023, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0103]

[0127] DSC

[0128] Conventional DSCs do not offer variable quality compression. Rather, DSCs take 24-bit color coding and compress it to 15 / 12 / 10 / 8 bits. The higher the compression (24→8bpp), the worse the impact on quality. With respect to the quality required for the section the eye is focusing on, embodiments can maintain PSNR quality settings above 60 dB, as discussed above. From the use case analysis shown in Figure 6, the inventors determined that this occurs only at a 37% compression configuration (24→15bpp). However, only the area the eye is currently focusing on actually utilizes that compression setting. The lateral foveal region (e.g., the part of the image further from the line of sight) can have lower quality, for example, a 75% compression level (24→8bpp).

[0104]

[0129] Therefore, in neighborhood-based compression standards such as DSC, which lack the concept of tiles, embodiments divide the main screen into a high-quality region and a low-quality region (as shown in Figure 18) or smaller sections (as shown in Figure 20) each having a different compression ratio. The selected compression ratio is a function of the current viewing position. Thus, referring to Figure 18, where the viewing position is located inside the high-quality region 1810, the high-quality region 1810 can be compressed at a lower compression level (e.g., 24 → 15 bpp), and the low-quality region 1820 can be compressed at a higher compression level (e.g., 24 → 8 bpp). In some examples, the low-quality region 1820 can be compressed at an even higher compression level (e.g., 24 → 6 bpp). In embodiments where the entire image is compressed using a higher compression level, as will be described more fully herein, the high-quality region 1810 can be superimposed on the entire image when the image is reconstructed.

[0105]

[0130] Figure 19 is a simplified flowchart illustrating a method 1900 for compressing an image using different compression ratios for high-quality and low-quality regions, according to one embodiment of the present invention. Method 1900 includes determining the user's line of sight (1910), generating a foveal map containing a first region of the image and a second region of the image (1912), and compressing the first region using a first compression ratio and the second region using a second compression ratio (1914).

[0106]

[0131] The image may be an image contained within a video stream. Determining the user's gaze position can be achieved by utilizing an eye-tracking system that provides the gaze position as a function of time. The foveal map defines the compression ratio to which portions of the image are compressed, which varies as a function of the position in the image relative to the gaze position, with areas(s) closer to the gaze position being compressed using a lower compression ratio and areas(s) further away from the gaze position being compressed using a higher compression ratio. In the example shown in Figure 18, the foveal map contains two regions, but the present invention is not limited to this particular implementation, and three or more regions can be defined. In some embodiments, the foveal map includes a first region of the image and a second region of the image. Method 1900 may be referred to as N-directional compression (e.g., DSC, VDC-X, or JPEG), where N refers to the number of regions determined for the image. For example, based on the gaze position, high-quality regions, medium-quality regions surrounding the high-quality regions, and low-quality regions can be determined for the image. The technique of Method 1900 can then be used as three-directional compression with different compression ratios for each region.

[0107]

[0132] Referring back to Figure 18, in some examples, the low-quality region 1820 may encompass the entire image, including the portion of the image within the high-quality region 1810 characterized by the gaze position. When decoding a compressed image (for example, for reconstruction for display to a user), it may be desirable to decode sections of the image in parallel. For an image divided into a high-quality region 1810 and a low-quality region 1820, as in Figure 18, the low-quality region 1820 may be considered the entire image. For example, in the case of a 2-kilopixel × 2-kilopixel image (4 megapixels total), the low-quality region 1820 may be the entire 4-megapixel image and may be compressed using a high compression level (e.g., 24 → 8 bpp). The high-quality region 1810 may be determined based on the current gaze position and may be, for example, a 1-kilopixel × 1-kilopixel region (1 megapixel total). The high-quality region 1810 can be compressed using a low compression level (e.g., 24 → 15 bpp). Therefore, two DSC decoders can be used to decode a compressed image. During image reconstruction, the decoded high-quality regions can be overlaid on the decoded low-quality regions.

[0108]

[0133] Figure 20 is a simplified image showing an image frame divided into high-quality and low-quality sections according to one embodiment of the present invention. As will be discussed in more detail below, the divided image frame shown in Figure 20 can be used to define a foveal map that defines the compression ratio at which different sections of the image are compressed, such that the compression ratio or other compression quality metric varies as a function of the position in the image relative to the gaze position. For example, sections closer to the gaze position can be compressed using a lower compression ratio, and sections further from the gaze position can be compressed using a higher compression ratio.

[0109]

[0134] Referring to Figure 20, four sections 2010, 2012, 2014, and 2016, including the high-quality region 2002 (i.e., the region corresponding to the current gaze position), are compressed at a lower compression level (e.g., 24 → 15 bpp), while the remaining sections, which may be referred to as peripheral sections or low-quality sections, are compressed at a higher compression level (e.g., 24 → 8 bpp). As a result, when the compressed image is reconstructed for display to the user, the high-quality region corresponding to the gaze position is characterized by higher quality than the rest of the image further from the gaze position. Consequently, embodiments of the present invention provide gaze-position-based fovealed images with reduced storage and transmission requirements.

[0110]

[0135] In some embodiments of the example shown in Figure 20, all sections 2010–2046 of the image may be compressed at a high compression ratio (e.g., 24 to 8 bpp). Four sections 2010, 2012, 2014, and 2016 containing high-quality regions may also be compressed at a lower compression ratio (e.g., 24 to 15 bpp). By using a decoder, all sections 2010–2046 compressed at a high compression ratio can be decoded according to a higher compression ratio, and the four sections 2010, 2012, 2014, and 2016 compressed at a lower compression ratio can be decoded according to a lower compression ratio. During image reconstruction, the decoded high-quality sections 2010, 2012, 2014, and 2016 can be superimposed on the decoded low-quality sections 2010–2046. In some embodiments, a foveal map can define sections that coincide with high-quality regions. For example, sections 2010–2016 may include only high-quality regions characterized by gaze position, excluding portions of images within low-quality regions.

[0111]

[0136] Similar to N-direction compression, it may be desirable to decode a compressed image using section-based DSC techniques with multiple DSC decoders. For example, a compressed image can be decoded using four DSC decoders: one decoder used to decode high-quality sections 2010-2016, another used to decode sections 2020-2026, a third used to decode sections 2030-2036, and a fourth used to decode sections 2040-2046, with each decoder using a compression ratio for each group of sections based on its proximity to the viewing position. In some embodiments, a single decoder can be implemented with acceptable latency when decoding a compressed image, depending on the memory capacity (e.g., SRAM) of the system used for decoding.

[0112]

[0137] The image may be an image contained within a video stream. Determining the user's gaze position can be achieved by utilizing an eye-tracking system that provides the gaze position as a function of time. The foveal map defines the compression ratios to which different sections of the image (e.g., sections 2010-2016, 2020-2026, 2030-2036, and 2040-2046) are compressed, which vary as a function of their position in the image relative to the gaze position, with sections closer to the gaze position being compressed using lower compression ratios and sections further away from the gaze position being compressed using higher compression ratios. In the example shown in Figure 20, the foveal map contains 16 sections, but the present invention is not limited to this particular implementation, and more or fewer sections may be defined. The methods described herein may be referred to as section-based compression (e.g., DSC, VDC-X, or JPEG) methods.

[0113]

[0138] While some of the examples above show only two compression levels, embodiments of the present invention are not limited to these specific compression levels, and an additional number of compression levels can be utilized. For example, sections 2010–2014 can be compressed using a 37% compression level (i.e., 24→15bpp), while sections 2020, 2022, 2024, and 2026, which are further from the high-quality region, can be compressed using a 50% compression level (i.e., 24→12bpp), sections 2030, 2032, 2034, and 2036, which are even further from the high-quality region than sections 2020–2026, can be compressed using a 58% compression level (i.e., 24→10bpp), and sections 2040, 2042, 2044, and 2046, which are the furthest from the high-quality region than sections 2010–2016, can be compressed using a 67% compression level (i.e., 24→8bpp). Therefore, the use of two compression levels is merely illustrative. Furthermore, for some sections, the compression level may be 0%, i.e., uncompressed, including sections corresponding to the viewing position and high-quality areas. Thus, a compressed image may have both uncompressed and compressed sections. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0114]

[0139] Furthermore, although Figure 20 shows only 16 sections of uniform area, this is not mandatory, and other numbers of sections with different sizes can be used, with smaller sections adjacent to the high-quality areas, and larger sections, such as those compressed at a higher level, being further away from the high-quality areas. Thus, the number of compression levels, the compression levels, the number of sections, and the size of the sections can vary depending on the specific application. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0115]

[0140] When image compression reduces the frame size, the communication interface, such as the MIPI interface, can be changed to enter a low-power data transmission mode, or even an ultra-low-power sleep mode, thereby saving computing resources and reducing power consumption. At the endpoint, the reconstruction of the compressed image can be performed before it is displayed to the user.

[0116]

[0141] Figure 21 is a simplified flowchart illustrating a method 2100 for compressing an image using different compression ratios for high-quality and low-quality sections, according to one embodiment of the present invention. Method 2100 includes determining the user's line of sight (2110), generating a foveal map containing a first section of the image and a second section of the image (2112), and compressing the first region using a first compression ratio and the second region using a second compression ratio (2114).

[0117]

[0142] It should be recognized that the specific steps shown in Figures 19 and 21 provide a specific method for compressing an image according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figures 19 and 21 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0118]

[0143] VDC-X

[0144] The VDC-X compression standard (e.g., VDC-M) uses a tile-based method instead of the nearest neighbor method. This compression standard encodes different tiles with different quality settings, but the objective of this conventional compression is to maintain a constant frame size (i.e., bitrate) overall. Therefore, the compression ratio, when selected, varies for each tile to maintain a constant bitrate. Using this compression standard in conjunction with embodiments of the present invention, video images are compressed based on the user's gaze position, not solely on the bitrate. As an example, four sections 2010, 2012, 2014, and 2016, which contain high-quality areas (i.e., areas corresponding to the current gaze position), are compressed with a higher quality setting than the remaining sections, which may be referred to as peripheral sections, and the remaining sections are compressed with a lower quality setting than those used for sections 2010-2016.

[0119]

[0145] In some embodiments of the present invention, since a constant bitrate is not maintained, the size of each frame changes over time, and the transport interface, such as MIPI, enters a low-power mode when not in use.

[0120]

[0146] Similar to the DSC-based method discussed above, in the VDC-X tile-based method, the embodiment encodes the quality of each tile based on the current position of the user's gaze. As shown in Figure 20, using gaze information provided by the AR system's gaze tracking system, the tiles are compressed using the VDC-X standard as a function of the distance of the tile from the gaze position.

[0121]

[0147] Therefore, embodiments of the present invention allow for variations in frame size or per-frame bitrate, and use the current line-of-sight information to select which tiles (VDC-X) or sections (DSC) have higher quality than foveal regions with lower quality settings.

[0122]

[0148] In some embodiments, the N-directional compression or section-based compression described above can implement JPEG as a compression standard rather than DSC or VDC-X. In these embodiments, the compression ratio used for high-quality / low-quality regions and / or high-quality / low-quality sections can instead refer to the quality settings of the JPEG standard.

[0123]

[0149] Figure 22 is a series of images illustrating the use of a spherical depth map according to one embodiment of the present invention. As shown in Figure 22, the images can be characterized by latitude and longitude frames. The latitude and longitude frames can be projected to form a 360° image. Using embodiments of the present invention, different portions of the 360° image can be generated from a desired viewpoint or perspective.

[0124]

[0150] Therefore, in the embodiment shown in Figure 22, instead of using a predicted head pose and corresponding content (e.g., an image or video stream), content captured by virtual or multiple cameras can be generated and projected to form latitude / longitude-based content or 360°-based content based on a given pose. This content can then be used with either the actual head pose or adjustments to perform reprojection and generate reprojected content corresponding to the actual head pose.

[0125]

[0151] Figure 23 is a simplified block diagram showing the components of an AR system according to one embodiment of the present invention. The AR system 2300 shown in Figure 23 can be incorporated into an AR device as described herein. Figure 23 provides a schematic diagram of one embodiment of the AR system 2300 that can carry out some or all of the steps of the method provided by various embodiments. It should be noted that Figure 23 is intended only to provide a generalized description of the various components, some or all of which may be appropriately utilized. Thus, Figure 23 broadly illustrates how individual system elements may be implemented in a relatively isolated or relatively more integrated manner.

[0126]

[0152] The AR system 2300 is shown as comprising hardware elements that can be electrically coupled via bus 2305 or otherwise communicate as needed. The hardware elements may include, but are not limited to, one or more processors 2310, including one or more general-purpose processors and / or one or more dedicated processors such as digital signal processing chips, graphics accelerators, and / or the like; one or more input devices 2315, which may include, but are not limited to, a mouse, keyboard, camera, and / or the like; and one or more output devices 2320, which may include, but are not limited to, a display device, printer, and / or the like. Additionally, the AR system 2300 includes an eye-tracking system 2355 that can provide the AR system with the user's gaze position. The foveal image compression techniques considered herein can be implemented using the processor 2310.

[0127]

[0153] The AR system 2300 may further include, but is not limited to, one or more non-temporary storage devices 2325 that may include local and / or network-accessible storage, and / or be able to communicate with them, and / or may include, but is not limited to, solid-state storage devices such as disk drives, drive arrays, optical storage devices, programmable, flash-updatable, and / or similar random-access memory (RAM) and / or read-only memory (ROM). Such storage devices may be configured to implement any suitable data store, including, but is not limited to, various file systems, database structures, and / or similar.

[0128]

[0154] The AR system 2300 may also include a communication subsystem 2319 which may include, but is not limited to, a chipset that may include a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a Bluetooth® device, an 802.11 device, a WiFi device, a WiMAX device, a cellular communication equipment, and / or similar. The communication subsystem 2319 may include one or more input and / or output communication interfaces to enable exchange of data with a network, such as, to give an example, another computer system, a television, and / or any other device described herein. Depending on the desired functionality and / or other implementation concerns, a portable electronic device or similar device may communicate images and / or other information via the communication subsystem 2319. In other embodiments, a portable electronic device, e.g., a first electronic device, may be incorporated into the AR system 2300, e.g., an electronic device as an input device 2315. In some embodiments, the AR system 2300 further includes a working memory 2360 which may include a RAM or ROM device, as described above.

[0129]

[0155] The AR system 2300 may also include software elements, shown as currently located in working memory 2360, including an operating system 2362, device drivers, executable libraries, and / or computer programs provided by various embodiments, and / or other code such as one or more application programs 2364 that may be designed to implement and / or configure the system as described herein, in accordance with the methods provided by other embodiments. As mere examples, one or more procedures described with respect to the methods considered above may be implemented as code and / or instructions executable by a computer and / or a processor within a computer. In one embodiment, such code and / or instructions may be used to configure and / or adapt a general-purpose computer or other device to perform one or more operations in accordance with the described methods.

[0130]

[0156] These instructions and / or sets of code can be stored in a non-temporary computer-readable storage medium such as the storage device(s) 2325 described above. In some cases, the storage medium may be incorporated into a computer system such as the AR system 2300. In other embodiments, the storage medium may be separate from the computer system, and may be a removable medium such as a compact disk, and / or provided in an installation package, so that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the stored instructions / code. These instructions may take the form of executable code that can be executed by the AR system 2300, and / or in the form of source and / or installable code, which takes the form of executable code when compiled and / or installed on the AR system 2300 using, for example, one of various commonly available compilers, installation programs, compression / decompression utilities, etc.

[0131]

[0157] It will be apparent to those skilled in the art that substantial modifications can be made according to specific requirements. For example, customized hardware may be used, and / or certain elements may be implemented in hardware, software including portable software such as applets, or both. Furthermore, connections to other computing devices, such as network input / output devices, may be used.

[0132]

[0158] As described above, in one embodiment, several embodiments can be used to implement methods according to various embodiments of the Art using a computer system such as the AR system 2300. According to a set of embodiments, some or all of the steps of such a method are implemented by the AR system 2300 in response to a processor 2310 that executes one or more sequences of one or more instructions that may be incorporated into an operating system 2362 and / or other code such as an application program 2364, which are contained in working memory 2360. Such instructions may be read into working memory 2360 from another computer-readable medium, such as one or more of the storage devices 2325. As just one example, the execution of a sequence of instructions contained in working memory 2360 can cause the processor 2310 to execute one or more steps of the method described herein. Additionally or alternatively, some of the methods described herein may be executed via dedicated hardware.

[0133]

[0159] As used herein, the terms machine-readable medium and computer-readable medium refer to any medium involved in providing data that causes a machine to operate in a particular way. In embodiments implemented using the AR system 2300, various computer-readable media may be involved in providing instructions / code to the processor(s) 2310 for execution, and / or may be used to store and / or carry such instructions / code. In many implementations, the computer-readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile medium or a volatile medium. Non-volatile media include, for example, optical and / or magnetic disks such as storage device(s) 2325. Volatile media include, but are not limited to, dynamic memory such as working memory 2360.

[0134]

[0160] Common forms of physical and / or tangible computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH-EPROMs, any other memory chips or cartridges, or any other media from which a computer can read instructions and / or code.

[0135]

[0161] Various forms of computer-readable media may be involved in transporting one or more sequences of one or more instructions to the processor(s) 2310 for execution. As just one example, the instructions may first be transported onto a magnetic disk and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions as signals over a transmission medium to be received and / or executed by the AR system 2300.

[0136]

[0162] The communication subsystem 2319 and / or its components generally receive signals, and the bus 2305 may then transport the signals and / or the data, instructions, etc. carried by the signals to the working memory 2360, from which the processor(s) 2310 retrieves and executes the instructions. Instructions received by the working memory 2360 may optionally be stored in a non-temporary storage device 2325 either before or after execution by the processor(s) 2310.

[0137]

[0163] Various embodiments of the present disclosure are provided below. When used herein, any reference to a set of embodiments should be understood as a disjunctive reference to each of those embodiments (for example, “Embodiments 1-4” should be understood as “Embodiments 1, 2, 3, or 4”).

[0138]

[0164] Example 1 is a method comprising: an encoder receiving virtual content; an encoder receiving a predicted head orientation corresponding to the virtual content; encoding the virtual content based on the predicted head orientation; generating compressed content; a decoder receiving the compressed content; a decoder receiving the current head orientation; decoding the compressed content based on the current head orientation; and generating reprojected virtual content.

[0139]

[0165] Example 2 is the method described in Example 1, wherein the virtual content corresponds to the predicted head pose, and the reprojected virtual content corresponds to the current head pose.

[0140]

[0166] Example 3 is the same as the method described in Examples 1-2, wherein the virtual content is an image of a set of multiple images.

[0141]

[0167] Example 4 is the method described in Examples 1 to 3, wherein the set of multiple images includes a right image and a left image.

[0142]

[0168] Example 5 is the same as the method described in Examples 1 to 4, wherein the color map and depth map are associated with virtual content.

[0143]

[0169] Example 6 is the method according to Examples 1-5, wherein encoding virtual content involves the use of a multi-view video coding process.

[0144]

[0170] Example 7 is the same as the method described in Examples 1 to 6, wherein the multiview video coding process is performed using an MV-HEVC encoder.

[0145]

[0171] Example 8 is the method according to Examples 1-7, wherein decrypting the compressed content involves using a multi-view video decoding process.

[0146]

[0172] Example 9 is the same as the method described in Examples 1 to 8, wherein the multiview video decoding process is performed using an MV-HEVC decoder.

[0147]

[0173] Example 10 is the method according to Examples 1 to 9, wherein the virtual content includes a video stream.

[0148]

[0174] Example 11 is the method described in Examples 1 to 10, wherein the current head posture is expressed as an adjustment to the predicted head posture.

[0149]

[0175] Example 12 is the method according to Examples 1 to 11, further comprising compressing the reprojected virtual content to form one or more foveated images, decompressing one or more foveated images to form a set of output images, and displaying the set of output images on a display.

[0150]

[0176] Example 13 is a method of Examples 1 to 12, wherein compressing the reprojected virtual content includes determining the user's gaze position and generating a foveal map based on the gaze position, wherein the foveal map includes a first region of the reprojected virtual content and a second region of the reprojected virtual content, and compressing the first region using a first quality setting and compressing the second region using a second quality setting.

[0151]

[0177] Example 14 is the same as the method described in Examples 1 to 13, wherein determining the gaze position involves the use of an eye-tracking camera of an augmented reality device.

[0152]

[0178] Example 15 is the method according to Examples 1 to 14, wherein the foveal map includes the central region and the peripheral region.

[0153]

[0179] Example 16 is a method of Examples 1 to 15, wherein compressing a first region using a first quality setting includes compressing all blocks within the first region using a first quality setting.

[0154]

[0180] Example 17 is the method described in Examples 1 to 16, wherein the first quality setting is higher than the second quality setting.

[0155]

[0181] Example 18 is the method described in Examples 1 to 17, wherein the first quality setting is 100%.

[0156]

[0182] Example 19 is a method of Examples 1 to 18, wherein decompressing one or more fovealized images includes decoding one or more fovealized images using a foveal map.

[0157]

[0183] Example 20 is the method according to Examples 1 to 19, wherein a first region of the image comprises a plurality of first blocks, a second region of the image comprises a plurality of second blocks, compressing the first region comprises compressing each of the plurality of first blocks using a first quality setting, and compressing the second region comprises compressing each of the plurality of second blocks using a second quality setting.

[0158]

[0184] Example 21 is the method according to Examples 1 to 20, wherein the second region includes the first region.

[0159]

[0185] Example 22 is the method according to Examples 1 to 21, wherein forming each of the set of output images includes decoding one or more fovealed images using a foveal map to generate a decoded first region and a decoded second region, and superimposing the decoded first region on top of the decoded second region.

[0160]

[0186] Example 19 is the method of Examples 1 to 22, wherein decompressing one or more fovealized images includes decoding one or more fovealized images using a foveal map.

[0161]

[0187] Embodiment 23 is a system comprising: a frame; one or more image capture devices coupled to the frame; a set of displays coupled to the frame; a set of projectors, each of which is optically coupled to one of the sets of displays; a memory; and a processor coupled to the memory, the processor configured to receive virtual content in an encoder, receive a predicted head pose corresponding to the virtual content in the encoder, encode the virtual content based on the predicted head pose, generate compressed content, receive the compressed content in a decoder, receive the current head pose in the decoder, decode the compressed content based on the current head pose, and generate virtual content to be reprojected.

[0162]

[0188] Example 24 is the system described in Example 23, wherein the virtual content corresponds to the predicted head pose, and the reprojected virtual content corresponds to the current head pose.

[0163]

[0189] Example 25 is the system described in Examples 23-24, wherein the virtual content is an image set of multiple images.

[0164]

[0190] Example 26 is the system described in Examples 23-25, wherein the set of multiple images includes a right image and a left image.

[0165]

[0191] Example 27 is a system described in Examples 23-26, in which a color map and a depth map are associated with virtual content.

[0166]

[0192] Example 28 is a system described in Examples 23-27, wherein encoding virtual content involves the use of a multi-view video coding process.

[0167]

[0193] Example 29 is a system described in Examples 23-28, wherein the multiview video coding process is performed using an MV-HEVC encoder.

[0168]

[0194] Example 30 is a system described in Examples 23-29, wherein decrypting compressed content involves the use of a multi-view video decoding process.

[0169]

[0195] Example 31 is a system described in Examples 23-30 in which the multiview video decoding process is performed using an MV-HEVC decoder.

[0170]

[0196] Example 32 is a system described in Examples 23-31, wherein the virtual content includes a video stream.

[0171]

[0197] Example 33 is a system described in Examples 23-32, in which the current head posture is represented as an adjustment to the predicted head posture.

[0172]

[0198] Example 34 is a system according to Examples 23-33, further configured in which the processor compresses the reprojected virtual content to form one or more foveated images, decompresses one or more foveated images to form a set of output images, and displays the set of output images on a display.

[0173]

[0199] Example 35 is a system according to Examples 23-34, wherein the compression of the reprojected virtual content includes determining the user's gaze position and generating a foveal map based on the gaze position, wherein the foveal map includes a first region of the reprojected virtual content and a second region of the reprojected virtual content, and compressing the first region using a first quality setting and compressing the second region using a second quality setting.

[0174]

[0200] Example 36 is a system described in Examples 23-35, wherein determining the gaze position involves the use of an eye-tracking camera of an augmented reality device.

[0175]

[0201] Example 37 is a system described in Examples 23 to 36, wherein the foveal map includes a central region and a peripheral region.

[0176]

[0202] Example 38 is a system according to Examples 23-37, wherein compressing a first region using a first quality setting includes compressing all blocks within the first region using a first quality setting.

[0177]

[0203] Example 39 is a system according to Examples 23 to 38, wherein a first region of the image comprises a plurality of first blocks, a second region of the image comprises a plurality of second blocks, compressing the first region comprises compressing each of the plurality of first blocks using a first quality setting, and compressing the second region comprises compressing each of the plurality of second blocks using a second quality setting.

[0178]

[0204] Example 40 is the system described in Examples 23 to 39, wherein the second region includes the first region.

[0179]

[0205] Example 41 is a system according to Examples 23-40, wherein forming each of the set of output images includes decoding one or more fovealed images using a foveal map to generate a decoded first region and a decoded second region, and superimposing the decoded first region on top of the decoded second region.

[0180]

[0206] Example 42 is a system described in Examples 23-41, wherein decompressing one or more fovealized images includes decoding one or more fovealized images using a foveal map.

[0181]

[0207] Example 43 is a system described in Examples 23 to 42, wherein the display set includes a right eyepiece waveguide display and a left eyepiece waveguide display.

[0182]

[0208] Embodiment 44 is a non-temporary computer-readable medium comprising program code executable by a processor of a user-wearable device, wherein the program code is executable by the processor to: receive virtual content in an encoder; receive a predicted head orientation corresponding to the virtual content in the encoder; encode the virtual content based on the predicted head orientation and generate compressed content; receive the compressed content in a decoder; receive the current head orientation in the decoder; decode the compressed content based on the current head orientation and generate reprojected virtual content.

[0183]

[0209] Example 45 is a non-temporary computer-readable medium according to Example 44, wherein the virtual content corresponds to the predicted head pose, and the reprojected virtual content corresponds to the current head pose.

[0184]

[0210] Example 46 is a non-temporary computer-readable medium as described in Examples 44-45, wherein the virtual content is an image of a set of multiple images.

[0185]

[0211] Example 47 is a non-temporary computer-readable medium described in Examples 44-46, wherein the set of multiple images includes a right image and a left image.

[0186]

[0212] Example 48 is a non-temporary computer-readable medium according to Examples 44-47, in which a color map and a depth map are associated with virtual content.

[0187]

[0213] Example 49 is a non-temporary computer-readable medium as described in Examples 44-48, wherein encoding the virtual content involves the use of a multi-view video coding process.

[0188]

[0214] Example 50 is a non-temporary computer-readable medium described in Examples 44-49, in which the multiview video coding process is performed using an MV-HEVC encoder.

[0189]

[0215] Example 51 is a non-temporary computer-readable medium as described in Examples 44-50, wherein decrypting the compressed content involves using a multi-view video decoding process.

[0190]

[0216] Example 52 is a non-temporary computer-readable medium according to Examples 44-51, in which the multiview video decoding process is performed using an MV-HEVC decoder.

[0191]

[0217] Example 53 is a non-temporary computer-readable medium as described in Examples 44-52, in which the virtual content includes a video stream.

[0192]

[0218] Example 54 is a non-temporary computer-readable medium described in Examples 44-53, in which the current head posture is represented as an adjustment to the predicted head posture.

[0193]

[0219] Example 55 is a non-temporary computer-readable medium according to Examples 44-54, wherein the processor is further configured to compress the reprojected virtual content to form one or more foveated images, decompress one or more foveated images to form a set of output images, and display the set of output images on a display.

[0194]

[0220] Example 56 is a non-temporary computer-readable medium according to Examples 44-55, wherein the compression of the reprojected virtual content includes determining the user's gaze position and generating a foveal map based on the gaze position, wherein the foveal map includes a first region of the reprojected virtual content and a second region of the reprojected virtual content, and compressing the first region using a first quality setting and compressing the second region using a second quality setting.

[0195]

[0221] Example 57 is a non-temporary computer-readable medium as described in Examples 44-56, in which determining the gaze position involves the use of an eye-tracking camera of an augmented reality device.

[0196]

[0222] Example 58 is a non-temporary computer-readable medium according to Examples 44-57, wherein the foveal map includes a central region and a peripheral region.

[0197]

[0223] Example 59 is a non-temporary computer-readable medium according to Examples 44-58, wherein compressing a first region using a first quality setting includes compressing all blocks within the first region using a first quality setting.

[0198]

[0224] Example 60 is a non-temporary computer-readable medium according to Examples 44-59, wherein a first region of the image comprises a plurality of first blocks, a second region of the image comprises a plurality of second blocks, compression of the first region comprises compressing each of the plurality of first blocks using a first quality setting, and compression of the second region comprises compressing each of the plurality of second blocks using a second quality setting.

[0199]

[0225] Example 61 is a non-temporary computer-readable medium according to Examples 44-60, wherein the second region includes the first region.

[0200]

[0226] Example 62 is a non-temporary computer-readable medium according to Examples 44-61, wherein forming each of the set of output images includes decoding one or more fovealed images using a foveal map to generate a decoded first region and a decoded second region, and superimposing the decoded first region on top of the decoded second region.

[0201]

[0227] Example 63 is a non-temporary computer-readable medium according to Examples 44-62, wherein decompressing one or more fovealized images includes decoding one or more fovealized images using a foveal map.

[0202]

[0228] In the aforementioned specification, this disclosure is described with reference to specific embodiments thereof. However, it will be apparent that various modifications and changes can be made without departing from the broader spirit and scope of this disclosure. Accordingly, this specification and the drawings should be considered illustrative rather than restrictive.

[0203]

[0229] In fact, it will be recognized that each of the systems and methods of this disclosure has several innovative aspects, and that not one alone alone can assume or be required for the desired attributes disclosed herein. The various features and processes described above may be used independently of each other or combined in various ways. All possible combinations and partial combinations are intended to fall within the scope of this disclosure.

[0204]

[0230] Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any preferred partial combination in multiple embodiments. Furthermore, features described above as acting in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be removed from the combination, and the claimed combination may cover a partial combination or a variation of a partial combination. Not a single feature or group of features is required or essential in every embodiment.

[0205]

[0231] In particular, conditional statements used herein, such as “can,” “could,” “might,” “may,” and “for example,” should generally be understood, unless otherwise specified or understood in the context in which they are used, to convey that a particular embodiment includes certain features, elements, and / or steps, but other embodiments do not. Therefore, such conditional statements are not generally intended to mean that features, elements, and / or steps are required in some way in one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are included in any particular embodiment, or should be implemented in a particular embodiment, with or without input or instruction from the author. Terms such as “comprising,” “including,” and “having” are synonymous and are used comprehensively and openly, without prejudice to additional elements, features, actions, behaviors, etc. Furthermore, the term "or" is used in its inclusive sense (rather than its exclusive sense), so for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In addition, as used in this application and the attached claims, the articles "a," "an," and "the" should be interpreted as meaning "one or more" or "at least one" unless otherwise specified. Similarly, while actions may be depicted in a particular order in the drawings, it should be recognized that, in order to achieve the desired result, such actions do not need to be performed in a particular order or sequential order shown, or that all shown actions will be performed. Moreover, the drawings may schematically illustrate one or more exemplary processes in the form of flowcharts. However, other actions not depicted may be incorporated into the schematicly illustrated exemplary methods and processes. For example, one or more additional actions may be performed before, after, simultaneously with, or in between any of the illustrated actions.Additionally, the operations may be configured or ordered in other embodiments. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other embodiments are within the scope of the following claims. In some cases, the operations described in the claims may be performed in a different order and still achieve the desired results.

[0206]

[0232] Accordingly, the claims are not intended to be limited to the embodiments shown herein, but should be given the broadest scope consistent with this disclosure, the principles and novel features disclosed herein. Accordingly, it should be understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or changes in light thereof should be suggested to those skilled in the art and should be included within the spirit and scope of this application and the appended claims.

Claims

1. It is a method, In an encoder, receiving virtual content and The encoder receives a predicted head pose corresponding to the virtual content, Encoding the virtual content based on the predicted head pose, Generating compressed content and In the decoder, receiving the compressed content and The decoder receives the current head posture, Decrypting the compressed content based on the current head posture, To generate virtual content that is reprojected, Methods that include...

2. The method according to claim 1, wherein the virtual content corresponds to the predicted head posture, and the reprojected virtual content corresponds to the current head posture.

3. The method according to claim 1, wherein the virtual content is an image of a set of multiple images.

4. The method according to claim 3, wherein the set of multiple images includes a right image and a left image.

5. The method according to claim 1, wherein a color map and a depth map are associated with the virtual content.

6. The method according to claim 1, wherein encoding the virtual content includes using a multiview video coding process.

7. The method according to claim 6, wherein the multiview video coding process is performed using an MV-HEVC encoder.

8. The method according to claim 1, wherein decoding the compressed content includes using a multiview video decoding process.

9. The method according to claim 8, wherein the multiview video decoding process is performed using an MV-HEVC decoder.

10. The method according to claim 1, wherein the virtual content includes a video stream.

11. The method according to claim 1, wherein the current head posture is expressed as an adjustment to the predicted head posture.

12. The reprojected virtual content is compressed to form one or more fovealed images, Decompressing one or more of the aforementioned foveal-shaped images to form a set of output images, The set of output images is displayed on the display, The method according to claim 1, further comprising:

13. Compressing the reprojected virtual content Determining the user's line of sight, The method involves generating a foveal map based on the gaze position, wherein the foveal map includes a first region of the reprojected virtual content and a second region of the reprojected virtual content. Compressing the first region using a first quality setting, and compressing the second region using a second quality setting, The method according to claim 12, including the method described in claim 12.

14. The method according to claim 13, wherein determining the gaze position includes using an eye-tracking camera of an augmented reality device.

15. The method according to claim 13, wherein the foveal map includes a central region and a peripheral region.

16. The method according to claim 13, wherein compressing the first region using the first quality setting includes compressing all blocks within the first region using the first quality setting.

17. The method according to claim 16, wherein the first quality setting is higher than the second quality setting.

18. The method according to claim 17, wherein the first quality setting is 100%.

19. The first region includes a plurality of first blocks, The second region includes a plurality of second blocks, Compressing the first region includes compressing each of the plurality of first blocks using the first quality setting, The method according to claim 13, wherein compressing the second region includes compressing each of the plurality of second blocks using the second quality setting.

20. The method according to claim 13, wherein the second region includes the first region.

21. To form each of the aforementioned sets of output images, Using the foveal map, one or more fovealed images are decoded to generate a decoded first region and a decoded second region. The decoded first region is superimposed on the decoded second region, The method according to claim 20, including the method described in claim 20.

22. The method according to claim 12, wherein decompressing the one or more fovealed images includes decoding the one or more fovealed images using the foveal map.

23. It is a system, Frame and, One or more image capture devices coupled to the frame, A set of displays coupled to the frame, A set of projectors, wherein each of the projectors in the set is optically coupled to one of the displays in the set, Memory and A processor coupled to the memory, wherein the processor In the encoder, virtual content is received, The encoder receives a predicted head pose corresponding to the virtual content, The virtual content is encoded based on the predicted head posture, Generate compressed content, In the decoder, the compressed content is received, The decoder receives the current head posture, Based on the current head posture, the compressed content is decoded. Generates virtual content that will be reprojected. A processor configured in such a way, A system equipped with these features.

24. The system according to claim 23, wherein the virtual content corresponds to the predicted head posture, and the reprojected virtual content corresponds to the current head posture.

25. The system according to claim 23, wherein the virtual content is an image of a set of multiple images.

26. The system according to claim 25, wherein the set of multiple images includes a right image and a left image.

27. The system according to claim 23, wherein a color map and a depth map are associated with the virtual content.

28. The system according to claim 23, wherein encoding the virtual content includes using a multiview video coding process.

29. The system according to claim 28, wherein the multiview video coding process is performed using an MV-HEVC encoder.

30. The system according to claim 23, wherein decrypting the compressed content includes using a multiview video decoding process.

31. The system according to claim 30, wherein the multiview video decoding process is performed using an MV-HEVC decoder.

32. The system according to claim 23, wherein the virtual content includes a video stream.

33. The system according to claim 23, wherein the current head posture is expressed as an adjustment to the predicted head posture.

34. The aforementioned processor, The reprojected virtual content is compressed to form one or more fovealed images. Decompress the one or more foveal-shaped images mentioned above to form a set of output images. The set of output images is displayed on the screen. The system according to claim 23, further configured as follows.

35. Compressing the reprojected virtual content Determining the user's line of sight, The process involves generating a foveal map based on the aforementioned gaze position, wherein the generated foveal map includes a first region of the reprojected virtual content and a second region of the reprojected virtual content. Compressing the first region using a first quality setting, and compressing the second region using a second quality setting, The system according to claim 34, including the system described in claim 34.

36. The system according to claim 35, wherein determining the gaze position includes using an eye-tracking camera of an augmented reality device.

37. The system according to claim 35, wherein the foveal map includes a central region and a peripheral region.

38. The system according to claim 35, wherein compressing the first region using the first quality setting includes compressing all blocks within the first region using the first quality setting.

39. The first region of the aforementioned image includes a plurality of first blocks, The second region of the aforementioned image includes a plurality of second blocks, Compressing the first region includes compressing each of the plurality of first blocks using the first quality setting, The system according to claim 35, wherein compressing the second region includes compressing each of the plurality of second blocks using the second quality setting.

40. The system according to claim 35, wherein the second region includes the first region.

41. To form each of the aforementioned sets of output images, Using the foveal map, one or more fovealed images are decoded to generate a decoded first region and a decoded second region. The decoded first region is superimposed on the decoded second region, The system according to claim 40, including the system described in claim 40.

42. The system according to claim 23, wherein decompressing the one or more fovealed images includes decoding the one or more fovealed images using the foveal map.

43. The system according to claim 23, wherein the set of displays includes a right eyepiece waveguide display and a left eyepiece waveguide display.

44. A non-temporary computer-readable medium comprising program code executable by the processor of a device that can be worn by a user, wherein the program code is executed by the processor, In the encoder, virtual content is received, The encoder receives a predicted head pose corresponding to the virtual content, The virtual content is encoded based on the predicted head posture, Generate compressed content, In the decoder, the compressed content is received, The decoder receives the current head posture, Based on the current head posture, the compressed content is decoded. Generates virtual content that will be reprojected. A non-temporary, computer-readable medium that is executable in such a way.

45. The non-temporary computer-readable medium according to claim 44, wherein the virtual content corresponds to the predicted head posture, and the reprojected virtual content corresponds to the current head posture.

46. The non-temporary computer-readable medium according to claim 44, wherein the virtual content is an image of a set of multiple images.

47. The non-temporary computer-readable medium according to claim 46, wherein the set of multiple images includes a right image and a left image.

48. The non-temporary computer-readable medium according to claim 44, wherein a color map and a depth map are associated with the virtual content.

49. The non-temporary computer-readable medium according to claim 44, wherein encoding the virtual content includes using a multiview video coding process.

50. The non-temporary computer-readable medium according to claim 49, wherein the multiview video coding process is performed using an MV-HEVC encoder.

51. The non-temporary computer-readable medium according to claim 44, wherein decrypting the compressed content includes using a multiview video decoding process.

52. The non-temporary computer-readable medium according to claim 51, wherein the multiview video decoding process is performed using an MV-HEVC decoder.

53. The non-temporary computer-readable medium according to claim 44, wherein the virtual content includes a video stream.

54. The non-temporary computer-readable medium according to claim 44, wherein the current head posture is represented as an adjustment to the predicted head posture.

55. The aforementioned processor, The reprojected virtual content is compressed to form one or more fovealed images. Decompress the one or more foveal-shaped images mentioned above to form a set of output images. The set of output images is displayed on the screen. A non-temporary computer-readable medium according to claim 44, further configured as follows.

56. Compressing the reprojected virtual content Determining the user's line of sight, The process involves generating a foveal map based on the aforementioned gaze position, wherein the generated foveal map includes a first region of the reprojected virtual content and a second region of the reprojected virtual content. Compressing the first region using a first quality setting, and compressing the second region using a second quality setting, A non-temporary computer-readable medium according to claim 55, including the following:

57. The non-temporary computer-readable medium according to claim 56, wherein determining the gaze position includes using an eye-tracking camera of an augmented reality device.

58. The non-temporary computer-readable medium according to claim 56, wherein compressing the first region using the first quality setting comprises compressing all blocks within the first region using the first quality setting.

59. The first region includes a plurality of first blocks, The second region includes a plurality of second blocks, Compressing the first region includes compressing each of the plurality of first blocks using the first quality setting, The non-temporary computer-readable medium according to claim 56, wherein compressing the second region comprises compressing each of the plurality of second blocks using the second quality setting.

60. The non-temporary computer-readable medium according to claim 56, wherein the second region includes the first region.

61. To form each of the aforementioned sets of output images, Using the foveal map, one or more fovealed images are decoded to generate a decoded first region and a decoded second region. The decoded first region is superimposed on the decoded second region, A non-temporary computer-readable medium according to claim 60, including the following:

62. The non-temporary computer-readable medium according to claim 56, wherein the foveal map includes a central region and a peripheral region.

63. The non-temporary computer-readable medium according to claim 55, wherein decompressing the one or more fovealed images includes decoding the one or more fovealed images using the foveal map.