Method and system for performing foveal-based image compression based on gaze direction

A foveated image compression method using gaze direction enhances augmented reality systems by selectively applying higher quality compression to areas of interest, addressing display inefficiencies and resource strain.

JP2026510991APending Publication Date: 2026-04-10MAGIC LEAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2024-03-19
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing augmented reality systems face challenges in efficiently presenting virtual content alongside real-world elements due to the complexity of the human visual system, leading to suboptimal display technologies that strain processing and memory resources.

Method used

Implementing a foveated image compression method that uses gaze direction to selectively apply higher quality compression to areas of interest while reducing quality in peripheral vision, utilizing waveguides and optical elements to enhance display systems.

Benefits of technology

This approach reduces processing and memory requirements while maintaining user experience by focusing resources on areas of gaze, thus optimizing display efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510991000001_ABST
    Figure 2026510991000001_ABST
Patent Text Reader

Abstract

The method for compressing an image includes determining the user's gaze position and generating a foveal map based on the gaze position. The foveal map includes a first region of the image and a second region of the image. The method also includes compressing the first region of the image using a first quality setting and compressing the second region of the image using a second quality setting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 453,376, entitled "Method and System for Performing Foveated Image Compression Based on Eye Gaze", filed on March 20, 2023, the disclosure of which is hereby incorporated by reference in its entirety for all purposes.

Background Art

[0002]

[0002] Modern computing and display technologies have facilitated the development of systems for so - called virtual reality or augmented reality experiences, and digitally reproduced images or portions thereof are presented to viewers in a way that they appear or can be perceived as if they were real. Virtual reality, i.e., VR scenarios, typically involve presenting digital or virtual image information without transparency to other actual real - world visual inputs, and augmented reality, i.e., AR scenarios, typically involve presenting digital or virtual image information as an extension to the visualization of the actual world around the viewer.

[0003]

[0003] Referring to FIG. 1, an augmented reality scene 100 is depicted. A user of AR technology views a setting such as a real - world park characterized by people, trees, buildings, and a concrete platform 120 in the background. The user also "sees" "virtual content" such as an image 110 of a robot standing on the real - world concrete platform 120 and a flying comic - like avatar character 102 that appears to be anthropomorphic of a honeybee. These elements 110 and 102 are "virtual" in that they do not exist in the real world. Due to the complexity of the human visual system, it is difficult to generate AR technology that facilitates a comfortable, natural, and rich presentation of virtual image elements surrounded by other virtual or real - world image elements.

[0004]

[0004] Despite these advances in display technology, there is a need in the art for improved methods and systems related to augmented reality systems, particularly display systems. [Overview of the Initiative]

[0005]

[0005] The present invention generally relates to methods and systems relating to projection display systems, including wearable displays. More specifically, embodiments of the present invention provide methods and systems that combine the concept of fovea (i.e., the degradation of video quality in sections that the human eye is not focusing on) with the concept of compression. The present invention is applicable to a variety of applications in computer vision and image display systems and light field projection systems, including stereoscopic systems and systems that deliver beamlets of light to the user's retina.

[0006]

[0006] The present invention achieves many advantages over conventional techniques. For example, embodiments of the present invention provide a method and system that allows portions of an image or video stream corresponding to the user's line of sight to be compressed using a higher quality setting than portions of the image or video stream further away from the user's line of sight. Thus, memory and processing resources can be saved while reducing or minimizing the impact on the user experience. These and other embodiments of the present invention, along with many of their advantages and features, will be described in more detail below in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0007] [Figure 1] This diagram shows a user's view of augmented reality (AR) through an AR device. [Figure 2A] These are cross-sectional side views of an example of a stacked set of waveguides, each containing an internally coupled optical element. [Figure 2B] Figure 2A is a perspective view of an example of one or more stacked waveguides. [Figure 2C]Figures 2A and 2B are top views of an example of one or more stacked waveguides. [Figure 3] This is a simplified diagram of an eyepiece waveguide with combined pupil expanders according to one embodiment of the present invention. [Figure 4] This figure shows an example of a wearable display system according to one embodiment of the present invention. [Figure 5] This is a perspective view of a wearable device according to one embodiment of the present invention. [Figure 6] This figure shows run-length coding of a quantized DCT block according to one embodiment of the present invention. [Figure 7] This is a diagram showing the JPEG header structure. [Figure 8] This is a diagram showing images compressed using a single quality setting. [Figure 9] This diagram shows a fovealed image having three fovealed regions according to one embodiment of the present invention. [Figure 10] This diagram shows a fovealed image with post-processing in the fovealed region according to another embodiment of the present invention. [Figure 11] This is a fovealed 3D generated image having three fovealed regions, according to yet another embodiment of the present invention. [Figure 12] This is a diagram showing an image that can be used with multiple foveal maps according to one embodiment of the present invention. [Figure 13] This is a simplified flowchart illustrating a method for compressing images using one embodiment of the present invention. [Figure 14] This is a simplified schematic diagram illustrating a gaze-based image foveal system according to one embodiment of the present invention. [Figure 15] The compression levels obtained as a function of time, expressed by consecutive frames versus frequency, for both a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention are shown. [Figure 16]The following shows a histogram of frame count versus compression for a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. [Figure 17] This is a simplified flowchart illustrating a method for compressing image frames using an alternating compression algorithm according to one embodiment of the present invention. [Figure 18] This is a simplified image showing an image frame divided into high-quality and low-quality regions according to one embodiment of the present invention. [Figure 19] This is a simplified flowchart illustrating a method for compressing an image using different compression ratios for high-quality and low-quality regions, according to one embodiment of the present invention. [Figure 20] This is a simplified image showing an image frame divided into high-quality tiles and low-quality tiles according to one embodiment of the present invention. [Figure 21] This is a simplified flowchart illustrating a method for compressing images using different compression ratios for high-quality and low-quality tiles according to one embodiment of the present invention. [Figure 22] This is a simplified block diagram showing the components of an AR system according to one embodiment of the present invention. [Modes for carrying out the invention]

[0008]

[0031] The drawings are referenced here, and the same reference numbers throughout refer to the same parts. Unless otherwise indicated, the drawings are schematic diagrams and are not necessarily drawn to scale.

[0009]

[0032] Referring now to FIG. 2A, in some embodiments, light impinging on the waveguide may need to be redirected so as to couple the light internally into the waveguide. An internal coupling optical element may be used to redirect and internally couple the light into its corresponding waveguide. Although referred to herein as an “internal coupling optical element,” the internal coupling optical element need not be an optical element and may be a non-optical element. FIG. 2A shows a cross-sectional side view of an example of a set 200 of stacked waveguides each including an internal coupling optical element. Each waveguide may be configured to output light of one or more different wavelengths, or one or more different wavelength ranges. Light from a projector is input into the set 200 of stacked waveguides and is externally coupled to the user as will be more fully described below.

[0010]

[0033] The illustrated set of stacked waveguides 200 includes waveguides 202, 204, and 206. Each waveguide includes associated internally coupled optical elements (which may also be referred to as optical input areas on the waveguide), for example, an internally coupled optical element 203 disposed on the main surface (e.g., upper main surface) of waveguide 202, an internally coupled optical element 205 disposed on the main surface (e.g., upper main surface) of waveguide 204, and an internally coupled optical element 207 disposed on the main surface (e.g., upper main surface) of waveguide 206. In some embodiments, one or more of the internally coupled optical elements 203, 205, and 207 may be disposed on the bottom main surface of each waveguide 202, 204, and 206 (particularly when one or more internally coupled optical elements are reflective deflection optical elements). As shown in the figures, the internally coupled optical elements 203, 205, and 207 may be located on the upper main surface of their respective waveguides 202, 204, and 206 (or on the upper part of the subsequent lower waveguide), particularly if those internally coupled optical elements are transmissive deflection optical elements. In some embodiments, the internally coupled optical elements 203, 205, and 207 may be located within the body of their respective waveguides 202, 204, and 206. In some embodiments, as considered herein, the internally coupled optical elements 203, 205, and 207 are wavelength-selective, selectively redirecting light of one or more wavelengths while transmitting light of other wavelengths. Although the internally coupled optical elements 203, 205, and 207 are shown on one side or corner of their respective waveguides 202, 204, and 206, it will be recognized that in some embodiments they may be located in other areas of their respective waveguides 202, 204, and 206.

[0011]

[0034] As shown in the figure, the internally coupled optical elements 203, 205, 207 may be offset horizontally from each other. In some embodiments, each internally coupled optical element may be offset such that light is received without passing through another internally coupled optical element. For example, each of the internally coupled optical elements 203, 205, 207 may be configured to receive light from different projectors and may be separated (e.g., horizontally spaced apart) from other internally coupled optical elements 203, 205, 207 so as to substantially not receive light from the other internally coupled optical elements 203, 205, 207.

[0012]

[0035] Each waveguide also includes associated optical distribution elements, such as optical distribution element 210 disposed on the main surface (e.g., upper main surface) of waveguide 202, optical distribution element 212 disposed on the main surface (e.g., upper main surface) of waveguide 204, and optical distribution element 214 disposed on the main surface (e.g., upper main surface) of waveguide 206. In some other embodiments, the optical distribution elements 210, 212, 214 may be disposed on the bottom main surfaces of the associated waveguides 202, 204, 206, respectively. In some other embodiments, the optical distribution elements 210, 212, 214 may be disposed on both the upper and bottom main surfaces of the associated waveguides 202, 204, 206, respectively. Alternatively, the optical distribution elements 210, 212, 214 may be disposed on different surfaces of the upper and bottom main surfaces of different associated waveguides 202, 204, 206, respectively.

[0013]

[0036] Waveguides 202, 204, and 206 may be separated and isolated by layers of material, for example, gas, liquid, and / or solid. For example, as shown, layer 208 may separate waveguides 202 and 204, and layer 209 may separate waveguides 204 and 206. In some embodiments, layers 208 and 209 are formed from low refractive index materials (i.e., materials with a lower refractive index than the materials forming directly adjacent waveguides among waveguides 202, 204, and 206). Preferably, the refractive index of the materials forming layers 208 and 209 is 0.05 or more, or 0.10 or less, lower than the refractive index of the materials forming waveguides 202, 204, and 206. Advantageously, layers 208 and 209 with lower refractive indices can function as cladding layers that facilitate total internal reflection (TIR) ​​of light passing through waveguides 202, 204, and 206 (e.g., TIR between the upper and lower main surfaces of each waveguide). In some embodiments, layers 208 and 209 are formed from air. Although not shown, it will be recognized that the upper and lower parts of the illustrated set of waveguides 200 may include directly adjacent cladding layers.

[0014]

[0037] Preferably, for ease of manufacture and other considerations, the materials forming waveguides 202, 204, and 206 are similar or identical, and the materials forming layers 208, 209 are similar or identical. In some embodiments, the materials forming waveguides 202, 204, and 206 may differ between one or more waveguides, and / or the materials forming layers 208, 209 may differ while still maintaining the various refractive index relationships described above.

[0015]

[0038] Continuing to refer to Figure 2A, rays 218, 219, and 220 are incident on waveguide set 200. It will be recognized that rays 218, 219, and 220 may also be introduced into waveguides 202, 204, and 206 by one or more projectors (not shown).

[0016]

[0039] In some embodiments, the light rays 218, 219, and 220 have different characteristics, such as different wavelengths or different wavelength ranges, which may correspond to different colors. Each internally coupled optical element 203, 205, and 207 deflects the incident light so that it propagates through waveguides 202, 204, and 206 respectively by TIR. In some embodiments, each of the internally coupled optical elements 203, 205, and 207 selectively deflects one or more specific wavelengths of light, while allowing other wavelengths to pass through the underlying waveguide and associated internally coupled optical elements.

[0017]

[0040] For example, the internally coupled optical element 203 may be configured to deflect a ray 218 having a first wavelength or wavelength range, and to transmit rays 219 and 220 having different second and third wavelengths or wavelength ranges, respectively. The transmitted ray 219 collides with and is deflected by an internally coupled optical element 205 configured to deflect light of the second wavelength or wavelength range. The ray 220 is deflected by an internally coupled optical element 207 configured to selectively deflect light of the third wavelength or wavelength range.

[0018]

[0041] Continuing to refer to Figure 2A, the deflected rays 218, 219, and 220 are deflected so that they propagate through their corresponding waveguides 202, 204, and 206. That is, the internal coupling optical elements 203, 205, and 207 of each waveguide deflect the light into their corresponding waveguides 202, 204, and 206, thereby internally coupling the light within their corresponding waveguides. The rays 218, 219, and 220 are deflected by TIR at an angle that causes the light to propagate through their respective waveguides 202, 204, and 206. The rays 218, 219, and 220 propagate through their respective waveguides 202, 204, and 206 by TIR until they collide with the corresponding optical distribution elements 210, 212, and 214 of the waveguides, where they are externally coupled to provide the externally coupled ray 216.

[0019]

[0042] Referring now to Figure 2B, a perspective view of an example of the stacked waveguides in Figure 2A is shown. As described above, the internally coupled rays 218, 219, and 220 are deflected by the internally coupled optical elements 203, 205, and 207, respectively, and then propagate through waveguides 202, 204, and 206 by TIR, respectively. The rays 218, 219, and 220 then collide with the optical distribution elements 210, 212, and 214, respectively. The optical distribution elements 210, 212, and 214 deflect the rays 218, 219, and 220 so that they propagate toward the externally coupled optical elements 222, 224, and 226, respectively.

[0020]

[0043] In some embodiments, the light distribution elements 210, 212, and 214 are orthogonal pupil expanders (OPEs). In some embodiments, the OPEs deflect or distribute light to externally coupled optics 222, 224, and 226, and in some embodiments, they can also increase the beam or spot size of this light as it propagates to the externally coupled optics. In some embodiments, the light distribution elements 210, 212, and 214 may be omitted, and the internally coupled optics 203, 205, and 207 may be configured to deflect light directly to the externally coupled optics 222, 224, and 226. For example, referring to Figure 2A, the light distribution elements 210, 212, and 214 may be replaced by externally coupled optics 222, 224, and 226, respectively. In some embodiments, the externally coupled optics 222, 224, and 226 are exit pupils (EPs) or exit pupil expanders (EPEs) that direct light towards the user's eye. It will be recognized that the OPE may be configured to increase the dimensions of the eyebox on at least one axis, and the EPE may be configured to increase the eyebox on an axis intersecting the axis of the OPE, for example, an orthogonal axis. For example, each OPE may be configured to redirect a portion of the light that hits the OPE to the EPE in the same waveguide, while allowing the rest of the light to continue propagating through the waveguide. Upon hitting the OPE again, another portion of the remaining light is redirected to the EPE, and the rest of that portion continues propagating through the waveguide, and so on. Similarly, upon hitting the EPE, a portion of the colliding light is redirected out of the waveguide toward the user, and the rest of that light continues propagating through the waveguide until it hits the EPE again, at which point another portion of the colliding light is redirected out of the waveguide, and so on. As a result, a single beam of internally coupled light can be "duplicated" each time a portion of its light is redirected by the OPE or EPE, thereby forming a field of cloned light beams. In some embodiments, the OPE and / or EPE may be configured to resize the light beam. In some embodiments, the functions of the light distribution elements 210, 212, and 214 and the externally coupled optical elements 222, 224, and 226 are combined in a combined pupil expander, as discussed in relation to Figure 2E.

[0021]

[0044] Therefore, referring to Figures 2A and 2B, in some embodiments, the set of waveguides 200 includes waveguides 202, 204, 206 for each component color, internally coupled optics 203, 205, 207, optical distribution elements (e.g., OPE) 210, 212, 214, and externally coupled optics (e.g., EP) 222, 224, 226. Waveguides 202, 204, 206 may be stacked with air gaps / cladding layers between them. The internally coupled optics 203, 205, 207 redirect or deflect the incident light into their waveguides (by different internally coupled optics receiving light of different wavelengths). The light then propagates within each waveguide 202, 204, 206 at an angle that yields a TIR. In the illustrated example, ray 218 (e.g., blue light) is deflected by the first internal coupling optical element 203, then bounces along the waveguide, interacting with the optical distribution element (e.g., OPE) 210 and then the external coupling optical element (e.g., EP) 222 in the manner described above. Rays 219 and 220 (e.g., green light and red light, respectively) pass through waveguide 202, with ray 219 colliding with the internal coupling optical element 205 and being deflected. Ray 219 then bounces along waveguide 204 via TIR, proceeding to its optical distribution element (e.g., OPE) 212, and then to the external coupling optical element (e.g., EP) 224. Finally, ray 220 (e.g., red light) passes through waveguide 206 and collides with the optical internal coupling optical element 207 of waveguide 206. The internal optical coupling element 207 deflects the light ray 220 so that after the light ray propagates to the optical distribution element (e.g., OPE) 214 by TIR, it propagates to the external coupling optical element (e.g., EP) 226 by TIR. The external coupling optical element 226 then finally externally couples the light ray 220 to the viewer, who also receives externally coupled light from the other waveguides 202 and 204.

[0022]

[0045] Figure 2C is a top view of an example of the stacked waveguides shown in Figures 2A and 2B. As shown, waveguides 202, 204, and 206 may be vertically aligned with their associated optical distribution elements 210, 212, and 214 and associated external coupling optics 222, 224, and 226. However, as discussed herein, the internal coupling optics 203, 205, and 207 are not vertically aligned. Rather, the internal coupling optics are preferably non-overlapping (e.g., laterally spaced as seen in the top or plan view). As further discussed herein, this non-overlapping spatial arrangement facilitates one-to-one input of light from different resources to different waveguides, thereby enabling the unique coupling of a particular light source to a particular waveguide. In some embodiments, arrangements including non-overlapping, spatially separated internal coupling optics may be referred to as a pupil-shifting system, where the internal coupling optics in these arrangements can correspond to sub-pupils.

[0023]

[0046] Figure 3 is a simplified diagram of an eyepiece waveguide having a combined pupil expander according to one embodiment of the present invention. In the example shown in Figure 3, the eyepiece 310 utilizes a combined OPE / EPE region in a one-sided configuration. Referring to Figure 3, the eyepiece 310 includes a substrate 320 on which an internally coupled optical element 322 and a combined OPE / EPE region 324, also referred to as a combined pupil expander (CPE), are provided. The incident ray 330 is internally coupled via the internally coupled optical element 322 and externally coupled as an output ray 332 via the combined OPE / EPE region 324.

[0024]

[0047] The combined OPE / EPE region 324 includes grids corresponding to both OPE and EPE that spatially overlap in the x and y directions. In some embodiments, the grids corresponding to both OPE and EPE are located on the same side of the substrate 320 such that the OPE grid is superimposed on the EPE grid, or the EPE grid is superimposed on the OPE grid (or both). In other embodiments, the OPE grid is located on the opposite side of the substrate 320 from the EPE grid such that the grids spatially overlap in the x and y directions but are separated from each other in the z direction (i.e., in different planes). Thus, the combined OPE / EPE region 324 can be implemented in either a one-sided or two-sided configuration.

[0025]

[0048] Figure 4 shows an example of a wearable display system 430 in which various waveguides and associated systems disclosed herein may be integrated. Continuing with reference to Figure 4, the display system 430 includes a display 432 and various mechanical and electronic modules and systems to support the functionality of the display 432. The display 432 is wearable by a user 440 (also referred to as the viewer) of the display system and can be coupled to a frame 434 configured to position the display 432 in front of the user 440's eyes. In some embodiments, the display 432 may be considered eyewear. In some embodiments, a speaker 436 is coupled to the frame 434 and configured to be positioned adjacent to the user 440's ear canal (in some embodiments, another speaker, not shown, may be optionally positioned adjacent to the user's other ear canal to provide stereo / shapeable acoustic control). The display system 430 may also include one or more microphones or other devices for detecting sound. In some embodiments, the microphone is configured to allow the user to provide input or commands to the system 430 (e.g., selection of voice menu commands, natural language questions, etc.) and / or to enable voice communication with other people (e.g., with other users of a similar display system). The microphone may further be configured as a peripheral sensor for collecting voice data (e.g., sounds from the user and / or the environment). In some embodiments, the display system 430 may further include one or more outward-oriented environmental sensors configured to detect objects, stimuli, people, animals, places, or other aspects of the world around the user. For example, the environmental sensors may include one or more cameras, which may be positioned outward, for example, to capture images similar to at least a portion of the user 440's normal field of view.In some embodiments, the display system may also be separate from the frame 434 and may include peripheral sensors that can be attached to the user 440's body (e.g., the user 440's head, torso, limbs, etc.). In some embodiments, the peripheral sensors may be configured to acquire data characterizing the user 440's physiological state. For example, the sensors may be electrodes.

[0026]

[0049] The display 432 is operably coupled to a local data processing module by a communication link such as a wired lead wire or a wireless connection, and the local data processing module can be mounted in various configurations, such as being fixedly mounted to a frame 434, fixedly mounted to a helmet or hat worn by the user, embedded in headphones, or otherwise detachably mounted to the user 440 (e.g., in a backpack configuration, in a belt-mounted configuration). Similarly, sensors can be operably coupled to the local processor and data module by a communication link such as a wired lead wire or a wireless connection. The local processing and data module may include a hardware processor and digital memory such as non-volatile memory (e.g., flash memory or a hard disk drive), both of which can be used to assist in data processing, caching, and storage. Optionally, the local processor and data module may include one or more central processing units (CPUs), graphics processing units (GPUs), dedicated processing hardware, etc. The data may include a) data captured from sensors such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, a gyroscope, and / or other sensors disclosed herein (which may, for example, be operably coupled to frame 434 or otherwise attached to user 440), and / or b) data acquired and / or processed using a remote processing module 452 and / or a remote data repository 454 (including data relating to virtual content), possibly for passage to display 432 after such processing or retrieval. The local processing and data module may be operably coupled to the remote processing and data module 450 by a communication link 438, such as via a wired or wireless link, and the remote processing and data module may include the remote processing module 452, the remote data repository 454, and a battery 460.The remote processing module 452 and the remote data repository 454 can be coupled to the remote processing and data module 450 by communication links 456 and 458 so that these remote modules are operablely coupled to each other and available as resources to the remote processing and data module 450. In some embodiments, the remote processing and data module 450 may include one or more of the following: an image capture device, a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope. In some other embodiments, one or more of these sensors may be mounted on the frame 434 or may be a standalone structure communicating with the remote processing and data module 450 by a wired or wireless communication path.

[0027]

[0050] Continuing to refer to Figure 4, in some embodiments, the remote processing and data module 450 may include one or more processors configured to analyze and process data and / or image information, such as one or more central processing units (CPUs), graphics processing units (GPUs), and dedicated processing hardware. In some embodiments, the remote data repository 454 may include digital data storage facilities that may be available via the Internet or other networking configurations in a “cloud” resource configuration. In some embodiments, the remote data repository 454 may include one or more remote servers that provide information, such as information for generating augmented reality content, to the local processing and data module and / or the remote processing and data module 450. In some embodiments, all data is stored, all calculations are performed on the local processing and data module, and fully autonomous use from the remote module is possible. Optionally, an external system including a CPU, GPU, etc. (e.g., one or more processors, one or more computer systems) may perform at least part of the processing (e.g., generating image information, processing data) and provide and receive information from the illustrated module, for example, via a wireless or wired connection.

[0028]

[0051] Figure 5 shows a perspective view of a wearable device 500 according to one embodiment of the present invention. The wearable device 500 includes a frame 502 configured to support one or more projectors 504 at various positions along a surface facing the interior of the frame 502, as shown in the figure. In some embodiments, the projectors 504 can be mounted near the temples 506. Alternatively, or in addition, another projector can be placed at position 508. Such a projector may include, for example, one or more liquid crystal on silicon (LCoS) modules, micro-LED displays, or fiber scanning devices, or may operate in conjunction with them. In some embodiments, light from projectors 504 or projectors placed at position 508 can be directed into an eyepiece 510 for display to the user's eye. A projector placed at position 512 can be somewhat smaller due to the proximity this brings to the waveguide system. The closer the position, the less light can be lost when the waveguide system directs the light from the projector to the eyepiece 510. In some embodiments, the projector at position 512 can be used in conjunction with projector 504 or a projector positioned at position 508. Although not depicted, in some embodiments, the projector can also be positioned below the eyepiece 510. The wearable device 500 is depicted including sensors 514 and 516. Sensors 514 and 516 may take the form of forward-facing and lateral-facing optical sensors configured to characterize the real-world environment surrounding the wearable device 500.

[0029]

[0052] Embodiments of the present invention utilize an eye-tracking system to determine the user's gaze position and to utilize the gaze position for an image compression process. Referring to Figure 5, an eye-tracking camera 505 is positioned on frame 502 and can be used to track the user's gaze position using a wearable device 500. In other embodiments, the gaze position is determined using other eye-tracking systems, and the eye-tracking camera 505 shown in Figure 5 is merely illustrative. Among other functions, as will be described more fully herein, the image compression process, internal communications, and display used to compress and decompress virtual content for storage in memory can be modified according to the gaze position. For example, portions of an image or video stream corresponding to a gaze position can be compressed using a higher-quality compression process compared to other portions of the image or video stream located further away from the gaze position. Since these further portions of the image or video stream are in the user's peripheral vision, any impact on the user experience resulting from reduced compression quality may be smaller than the benefits achieved in terms of memory and processing efficiency and / or requirements. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0030]

[0053] Conventional systems implement fixed-quality image compression (e.g., JPEG compression) for images or video streams that do not take human gaze into consideration. Since MPEG is a derivative of JPEG, embodiments of the present invention can be appropriately applied to MPEG compression processing. By knowing where the human gaze is currently located and taking the human gaze into consideration, embodiments of the present invention can reduce the quality (i.e., bandwidth) in locations in the image that the user is not looking at, i.e., locations in the image that are spatially distant from the gaze position, thereby degrading the image quality in these areas and reducing the overall need to transmit something that the human eye would not be able to identify because the human eye is not currently focused on these non-gaze locations, at a high quality setting. Accordingly, embodiments of the present invention provide a video compression algorithm that takes human gaze into consideration and creates a foveal compression algorithm that is dependent on human gaze.

[0031]

[0054] In some embodiments, the JPEG algorithm receives an image and segments it into macroblocks (e.g., 16x16 pixels). These macroblocks then undergo a Discrete Cosine Transform (DCT) process. The DCT process generates a set of coefficients that are filtered so that high-frequency values ​​are removed (this is where the quality step comes in). After this process, the blocks are then run-length encoded.

[0032]

[0055] Encoder-based foveal map

[0056] Table 1 is a matrix showing an 8x8 pixel subimage block according to one embodiment of the present invention. An 8x8 pixel subimage block can also be called a macroblock or tile. The 8x8 pixels are represented by the pixel values ​​shown in the matrix. [Table 1]

[0033]

[0057] Table 2 is a matrix showing an example of an encoded 8x8 FDCT block according to one embodiment of the present invention. In conventional systems, JPEG / MPEG compression processes the entire image at a fixed quality. As a result of the filtering process, zero data is generated as shown in the quantized DCT block shown in Table 3. This filter occurs at a given quality setting. As shown in Table 2, the magnitude of the values ​​generally decreases from the upper left portion of the matrix to the lower right portion of the matrix. [Table 2]

[0034]

[0058] Table 3 is a matrix showing an example of a quantized DCT block according to one embodiment of the present invention. In Table 3, as a result of quantization, a considerable number of values ​​have been reduced to zero. [Table 3]

[0035]

[0059] Figure 6 shows run-length coding of a quantized DCT block according to one embodiment of the present invention. To encode the quantized DCT block, the run-length coding process begins with the top-left pixel and proceeds to the bottom-right pixel. Referring to Figure 6, pixel 610 is coded, followed by pixel 612. Next, the coding proceeds to the next two rows of pixels, resulting in the coding of pixels 614 and 616. The subsequent coding process results in the coding of pixels 618, 620, and 622. At this stage, the coding process reverses direction and codes pixels 624, 626, and 628.

[0036]

[0060] This encoding pattern then continues until all pixels within the block have been encoded.

[0037]

[0061] Figure 7 shows the JPEG header structure. As shown in Figure 7, in the JPEG header structure, the quantization table map area stores the default quality for the entire image. Therefore, a single quality setting is used to compress the entire image. As described herein, the quantization table can be applied to non-foveated regions to provide a 100% quality setting to the region corresponding to the line of sight, or the quantization table can be applied to regions further away from the line of sight to reduce the quality setting for foveated regions. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0038]

[0062] Referring to Figure 7, the segment includes the start of the image, application 0 (default header), definition of the quantization table (for luminance), defined quantization table (for chrominance), start of the frame, definition of Huffman table 1, definition of Huffman table 2, definition of Huffman table 3, definition of Huffman table 4, start of the scan, image data (entropy-encoded segment), and end of the image. The fields and values ​​of these segments are shown in Table 4. [Table 4] JPEG2026510991000006.jpg95170

[0039]

[0063] Embodiments of the present invention maintain high quality in blocks of image that the eye is focused on, while reducing the quality setting of blocks of image that the eye is not focusing on. These different quality settings are stored in a foveal map. The foveal map can then be passed to a compression engine. The compression engine can then selectively modify certain video blocks corresponding to the gaze position to compress these predetermined video blocks in high quality, while compressing other blocks in low quality.

[0040]

[0064] A foveal map can be created based on gaze information, that is, by actively communicating where the human eye is currently focused or looking. In embodiments of the present invention, the foveal map is supplied to an encoder and passed to a decoder.

[0041]

[0065] An additional advantage provided by embodiments of the present invention is that, by using the concept of video blocks, blocks with zero data (i.e., all black) reduce the memory space or power consumed during the video display process. Accordingly, embodiments of the present invention utilize a video block compression algorithm modified to implement variable quality on a per-block basis.

[0042]

[0066] Decoder-based foveal map

[0067] The decoder can use the current stream of DCT coefficients included as part of the compression standard, which has been passed to it. Therefore, some blocks will have more coefficients, and some will have fewer. However, a foveal map can be sent to or passed to the decoder so that it can use the location of the reduced-quality block / tile positions. Thus, the foveal map is used by the decoder to apply the desired quality setting to each tile / block. Additionally, this information can be used to apply post-processing image filtering to remove JPEG low-quality artifacts.

[0043]

[0068] Map implementation

[0069] It should be noted that certain implementations may have 100% inferred quality and may utilize a global table as an alternative table, or vice versa. Embodiments of the present invention may utilize various mechanisms for implementing quality map selection. As described herein, embodiments utilize two or more quality settings per image, and the quality settings are defined per tile / block. Thus, a foveal map supplied to an encoder (e.g., a JPEG encoder) allows the encoder to determine which quality setting is used for a given tile / block.

[0044]

[0070] Instead of two maps, you can also use three or more maps. The foveal index (0, 1, 2...) for each block indicates to the encoder which map to implement. Therefore, you can have quality settings in ranges such as 100%, 75%, 50%, 25%, etc.

[0045]

[0071] Figure 8 is a diagram showing an image compressed using a single quality setting. In this case, all pixels in the image are compressed using a conventional process that utilizes a single quality setting for each pixel. While this process achieves uniform image compression across the image, the inventors determined that processing and memory requirements can be reduced if portions of the image farther from the user's viewing position are compressed at a reduced quality compared to the portion of the image corresponding to the user's viewing position, while still achieving the desired user experience.

[0046]

[0072] Figure 9 is a diagram showing a fovealed image having three fovealed regions according to one embodiment of the present invention. The image in Figure 9 is divided into multiple regions based on the gaze position. In this case, the user is fixated on the center of the image, and as a result, the gaze position is located at the center of the image. As discussed herein, the gaze position can be determined using an eye-tracking system such as those discussed in relation to Figures 5 and 22. Thus, the image can be divided into a central region corresponding to the gaze position and peripheral regions further away from the gaze position. In some embodiments, a foveal map is created based on the gaze position, with portions of the image closer to the gaze position mapped to a high-quality setting and portions of the image further away from the gaze position mapped to a low-quality setting. In Figure 9, the foveal map takes the form of two peripheral regions with lower quality settings and a central region with a higher (e.g., 100%) quality setting.

[0047]

[0073] In the image shown in Figure 9, region 910, corresponding to the left quarter of the image (i.e., the left 1 / 4), is compressed using a first quality setting. Additionally, region 930, corresponding to the right quarter of the image (i.e., the right 1 / 4), is compressed using a first quality setting. However, region 920, corresponding to the middle half of the image (i.e., the central 2 / 4), is compressed using a second quality setting, which is higher than the first quality setting. This division of the image into parts can be referred to as a three-region division: the left quarter (e.g., fovealed with a 70% quality setting), the middle half (e.g., non-fovealed with a 100% quality setting), and the right quarter (e.g., fovealed with a 70% quality setting).

[0048]

[0074] Figure 9 shows the division into three regions using a foveal map containing these three regions, but the present invention is not limited to this implementation, and images can be divided in other ways. By dividing an image into multiple regions, the quality setting of individual blocks or tiles (e.g., 8x8 pixel blocks for JPEG compression) contained in each region can be set to a predetermined quality setting for each block. Thus, in Figure 9, the same quality setting is assigned to all blocks within each region, i.e., the blocks in region 910 are assigned a first quality setting (e.g., 70%), the blocks in region 920 are assigned a second quality setting (e.g., 100%), and the blocks in region 930 are assigned a first quality setting (e.g., 70%), but this is not mandatory, and different quality settings can be assigned to individual blocks within a region. Thus, the foveal map can be more complex than the three-region division shown in Figure 9. In some embodiments, blocks in the peripheral regions are assigned a quality setting that depends on the distance of the block from the line of sight, while blocks in the central region have a uniform quality setting. In other embodiments, the foveal map can be defined such that blocks in the peripheral region are assigned a uniform quality setting, while blocks in the central region are assigned a quality setting that depends on the distance of the block from the line of sight. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0049]

[0075] In the three-region fovealization image shown in Figure 9, an overall reduction of approximately 67% in image / memory size was achieved while maintaining 100% quality in region 920, i.e., the non-fovealized section. As discussed above, the non-fovealized region (i.e., compressed using an uncompressed or lossless compression algorithm) can be any region identified in the foveal map. Consequently, the three-region division shown in Figure 9 is merely illustrative.

[0050]

[0076] It should be noted that if the gaze position is, for example, on the right side of the image, the foveal map can compress the right side using a higher quality setting and the left side of the image using a lower quality setting. Therefore, in this example, if the gaze position is within region 930, regions 910 and 920 are compressed using a first quality setting, and region 930 is compressed using a second quality setting higher than the first. In some embodiments, for example, if the gaze position is within region 930, region 930 can be compressed using a higher quality setting, e.g., lossless compression, region 920 can be compressed using an intermediate quality setting lower than the higher quality setting, and region 910 can be compressed using a minimum quality setting lower than the intermediate quality setting. As a result, the fovea of ​​the image is a function of the gaze position, compressing or encoding the region containing the gaze position with a higher quality setting than one or more regions further away from the gaze position. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0051]

[0077] Furthermore, although Figure 9 shows a set of vertical regions, this is not essential to embodiments of the present invention, and the definition of regions can be carried out in other ways, including horizontally oriented regions, regions defined based on the distance to the line of sight, for example, a set of regions defined radially.

[0052]

[0078] Figure 10 shows a second foveated image with post-processing in the foveated region, according to another embodiment of the present invention. After post-processing of the image shown in Figure 9, the blurring of the image content in the foveated region, i.e., regions 910 and 930, reduces artifacts present in these regions.

[0053]

[0079] Figure 11 is a foveated 3D generated image having three foveated regions according to yet another embodiment of the present invention. In Figure 11, the regions are defined similarly to those shown in Figures 9 and 10. However, since the majority of the image in the 3D generated image is black, much higher compression is possible. Using the method described herein, an 87% compression was achieved while maintaining 100% quality at the center of the image corresponding to the gaze position. In this example, region 1120 was compressed using a 100% quality setting (non-foveated at 100% quality setting), while regions 1110 and 1130 were compressed using a lower quality setting (foveated at 20% quality setting). Since in many examples of virtual content the image content is highest near the gaze position and the surrounding areas are dark or black, embodiments of the present invention are particularly well suited for use in virtual reality and augmented reality implementations.

[0054]

[0080] In some examples, all areas of an image can be compressed using a lower quality setting, while non-foveated areas can be compressed using a higher quality setting. Using the example in Figure 9, areas 910, 920, and 930 can be compressed using a lower quality setting for the foveated areas, respectively. Area 920 can also be compressed using a higher quality setting. When decoding a compressed image (for example, for reconstruction for display to a user), it may be desirable to decode sections of the image in parallel. Therefore, two decoders can be used to decode a compressed image. During image reconstruction, the decoded area 920 using the higher quality setting can be superimposed on the decoded areas 910, 920, and 930 (i.e., the entire image) using the lower quality setting. The encoding may be JPEG (for example, using the quality settings described above), or it may be a technique including DSC or VDC-X (for example, using compression ratios), which are discussed more fully herein.

[0055]

[0081] Figure 12 is a diagram showing an image that can be used with multiple foveal maps according to one embodiment of the present invention. Figure 12 shows an image that includes a person 1206 located in section 1210, a tree 1202 located in sections 1220, 1222, 1230, and 1232, and a house 1204 located in sections 1224, 1226, 1238, and 1240. Different foveal maps can be created based on this image depending on the line of sight.

[0056]

[0082] If the user's gaze position is located in one of sections 1220, 1222, 1230, or 1232, i.e., if the user is looking at tree 1202, a foveal map can be used, and blocks of sections 1220, 1222, 1230, and 1232 are compressed using a 100% quality setting (non-fovealed at 100% quality), while blocks of the remaining sections (i.e., sections 1210, 1212, 1214, 1216, 1224, 1226, 1228, 1234, 1236, 1238, 1240, and 1242) are compressed using a lower quality setting (fovealed at 70% quality). Thus, image compression can be implemented using a foveal map that maintains quality within the region of the image corresponding to the gaze position, while peripheral parts of the image can be compressed using a lower quality setting to save system resources, including memory and processing.

[0057]

[0083] Alternatively, if the user's gaze position is in one of sections 1224, 1226, 1238, or 1240, i.e., if the user is looking at house 1204, the foveal map can be utilized, and the blocks in sections 1224, 1226, 1238, and 1240 are compressed using a 100% quality setting (non-fovealed at 100% quality), while the blocks in the remaining sections (i.e., sections 1210, 1212, 1214, 1216, 1220, 1222, 1228, 1230, 1232, 1234, and 1236, as well as 1242) are compressed using a lower quality setting (fovealed at 70% quality).

[0058]

[0084] Finally, when the user's gaze position is in section 1210, i.e., when the user is looking at person 1206, the foveal map can be utilized, and the blocks in section 1210 are compressed using a 100% quality setting (non-fovealed at 100% quality), while the blocks in the remaining sections (i.e., sections 1212, 1214, 1216, 1220, 1222, 1224, 1226, 1228, 1230, 1232, 1234, and 1236, 1238, 1240, and 1242) are compressed using a lower quality setting (fovealed at 70% quality). In some embodiments, the quality setting used for the remaining sections varies, for example, as a function of the distance from the gaze position. In these embodiments, the blocks in sections 1212, 1214, and 1216 can be compressed using a 90% quality setting, the blocks in sections 1220, 1222, 1224, 1226, and 1228 can be compressed using an 80% quality setting, and the blocks in sections 1230, 1232, 1234, and 1236, 1238, 1240, and 1242 can be compressed using a 70% quality setting. In some examples, instead of encoding with JPEG (e.g., using the quality settings described above), sections 1210-1242 may be compressed using techniques including DSC or VDC-X (e.g., using a compression ratio). For example, based on the gaze position, a non-tile-based compression technique such as DSC can be used to compress sections closer to the gaze position at a lower compression ratio, while sections further away from the gaze position can be compressed at a higher compression ratio.

[0059]

[0085] Figure 13 is a simplified flowchart illustrating a method for compressing an image according to one embodiment of the present invention. Method 1300 includes receiving an image (1310), determining the user's gaze position (1312), and generating a foveal map based on the gaze position (1314).

[0060]

[0086] The image may be an image contained within a video stream. Determining the user's gaze position can be achieved by utilizing an eye-tracking system that provides the gaze position as a function of time. The foveal map defines the quality at which blocks are compressed, which varies as a function of their position in the image; blocks in regions closer to the gaze position are compressed using a higher quality setting, and blocks in regions further away from the gaze position are compressed using a lower quality setting. In the example shown in Figure 9, the foveal map includes three regions, but the present invention is not limited to this particular implementation, and two or more regions can be defined. Furthermore, blocks within a given region can be compressed using a uniform quality setting, or they can be compressed with different quality settings depending on the particular implementation. In some embodiments, the foveal map includes a first region of the image and a second region of the image.

[0061]

[0087] The method also includes compressing a first region of an image using a first quality setting and compressing a second region of the image using a second quality setting (1316). In some embodiments, the first quality setting is an uncompressed quality setting or a lossless compression quality setting. Thus, blocks within the first region are compressed with a higher quality than other parts of the image. The second quality setting is a lower quality setting, e.g., a 70% quality setting, which reduces the data corresponding to the compressed image within these regions. As discussed above, since the user's line of sight places these regions in the user's peripheral vision, any loss of quality is offset by savings in memory and processor usage. The data compression processes for the first and second regions can be performed sequentially or in parallel, depending on the specific application.

[0062]

[0088] A compressed image or video, which may be called a foveal image or video, can be transmitted to a display system along with the foveal map (1318), or stored in memory along with the foveal map (1319).

[0063]

[0089] In embodiments in which a compressed image or video is stored in memory together with a foveal map, method 1300 includes retrieving the fovealized image and foveal map from memory (1320), decompressing a first region of the image using a first quality setting, and decompressing a second region of the image using a second quality setting (1340). In embodiments in which a compressed image or video is transmitted to a display system together with a foveal map, method 1300 includes receiving the fovealized image and foveal map (1320), decompressing a first region of the image using a first quality setting, and decompressing a second region of the image using a second quality setting (1340). The decompression processes for the first and second regions can be performed sequentially or in parallel, depending on the specific application. The two regions can be merged to form a final image suitable for display (1342). The final image is then displayed on a display device (1344).

[0064]

[0090] It should be noted that the specific steps shown in Figure 13 provide a specific method for compressing an image according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figure 13 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0065]

[0091] Figure 14 is a simplified schematic diagram showing a gaze-based image foveal system according to one embodiment of the present invention. Referring to Figure 14, the gaze-based image foveal system 1400 includes a wearable 1410 (e.g., a wearable including an ASIC that performs the illustrated operation) that receives an image or video suitable for display to the user. The image or video can be received using one or more communication interfaces 1420. In the illustrated embodiment, the image or video content is received using WiFi, USB, DisplayPort (DP) or other communication protocols. In this embodiment, the uncompressed content is MPEG video.

[0066]

[0092] The wearable 1410 also receives gaze information from the gaze tracking system 1405. The gaze tracking system 1405 may include one or more sensors suitable for measuring the position and orientation of the eyes and can provide data that can be used by the gaze processor 1430 when calculating the user's gaze. In the embodiment shown in Figure 14, the gaze processor 1430 is implemented using a CPU or a neural processing unit (NPU) controller, but other processors may also be used. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0067]

[0093] As shown in Figure 14, the image or video is passed to an image compression processor 1422 in some embodiments, which implements a process to form a compressed image / video (e.g., a fovealized image / video) based on the user's gaze, as will be discussed more fully herein. Different foveal processes, including tile-based foveal processes such as JPEG or DSC foveal processes, sparse-based compression processes, etc., as will be discussed more fully herein, can be appropriately utilized for specific applications. In some embodiments, the image compression processor 1422 is bypassed, for example, when the image is remotely compressed before it is received by one or more communication interfaces 1420 and the image or video is passed to memory 1424 for storage.

[0068]

[0094] When an image or video compressed using the image compression processor 1422, or remotely compressed, is retrieved from memory 1424, the image decompression process can be performed using gaze information provided by the decompression processor 1426 and the gaze processor 1430. In embodiments where the image is remotely compressed and the image compression processor 1422 is bypassed, the decompression processor 1426 can decode the compressed image. The original or reconstructed image is then passed to the warp / depth reprojection processor 1428.

[0069]

[0095] After warping or depth reprojection, the warped image can be compressed using a variable-quality encoder 1432, which includes a processor component 1431 representing the fovea of ​​the image based on the gaze position, by again utilizing the data provided by the gaze-line processor 1430. Different foveal processes, including tile-based foveal processes and sparse-based compression processes, can be appropriately utilized for specific applications. In some embodiments, the variable-quality encoder 1432, which includes the processor component 1431, is bypassed. As discussed above, the JPEG encoding process can be carried out by the variable-quality encoder 1432 to form a gaze-line based fovealed image in which the image quality varies across the image, providing high quality in the region of the image corresponding to the user's gaze position and lower quality in the region of the image further away from the gaze position. Thus, fovealed encoded images and sparse encoded images can be formed at a reduced size while maintaining the desired image quality. The encoded image is then provided to a Mobile Interface Processor Interface (MIPI) device 1434 for subsequent transmission to a display system.

[0070]

[0096] The MIPI device 1434 of the wearable 1410 can be connected to the MIPI device 1442 of a display system 1440, which includes a variable-quality decoder 1444 that includes a processor component 1443 that performs foveal desorption based on gaze position, and a display device 1446 such as an LCOS display or a micro-light-emitting diode (μLED) display. As shown in the implementation of the variable-quality decoder 1444 shown in Figure 14, JPEG / DSC tile-based encoded data or N-directional compression-based encoded data (e.g., N-directional DSC) can be received on a first communication channel, and a quality map (Q-map), such as a foveal map, can be received on a second communication channel for use during the decoding process. Alternatively, the Q-map can be received using an embedded line format or other preferred format.

[0071]

[0097] As shown in Figure 14, the JPEG decoding process is carried out by a variable-quality decoder 1444 including a processor component 1443, and the final image can be formed based on the foveated image generated by a variable-quality encoder 1432 including a processor component 1431. Thus, embodiments of the present invention reduce system memory and transmission requirements, for example, the amount of data transmitted between MIPI devices, while maintaining the desired image quality. The decoded image is then displayed using a display device 1446.

[0072]

[0098] In some embodiments, the variable-quality encoder 1432 is bypassed, and the warp image is transmitted to the display system 1440 using the MIPI device 1434 without variable-quality image compression. In these embodiments, the variable-quality decoder 1444 is also bypassed.

[0073]

[0099] Although the embodiments described above utilize a tile-based (also called block-based) JPEG compression algorithm, embodiments of the present invention are not limited to this particular compression standard, and other compression standards can be used in conjunction with various embodiments of the present invention. As an example, Figures 15 to 21 describe a technique that uses run-length coding with DSC and VDC-X to compress video data.

[0074]

[0100] Figure 15 shows the compression levels obtained as a function of time, expressed by consecutive frames versus frequency, for both a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. In Figure 15, each frame was compressed using either a mask-based compression method or DSC, according to alternating algorithms that implement either a mask-based compression method or fully fixed-frame compression, such as DSC.

[0075]

[0101] As shown in Figure 15, each frame is analyzed to determine the number of lines with pixels that have a brightness level below a threshold. If the mask-based compression method results in a compression level greater than the compression threshold (e.g., 37%), the frame is compressed using the mask-based compression method. In Figure 15, this results in the first approximately 3800 frames being compressed using the mask-based compression method.

[0076]

[0102] The DSC method is used when a mask-based compression method produces compressed frames with a compression level of less than 37%, such as frames with little black content. This results in these frames having a compression value of 37%. Referring to Figure 15, frames represented by blue compression values ​​of less than 37% are compressed using DSC, effectively baseline the minimum compression at 37%. Therefore, the frames in sets A and B have a compression value of 37%, rather than the lower values ​​achieved using the mask-based compression method.

[0077]

[0103] Figure 16 shows a histogram of frame count versus compression for a sparse compression system implementation and a DSC-SPARSE system implementation according to one embodiment of the present invention. As shown in Figure 16, the number of frames with less than approximately 37% compression is reduced to zero because the mask-based compression method was used for frames that could be compressed to a compression level greater than 37%, or the frame-based compression method (e.g., DSC) was used for the remaining frames that could not be compressed to a compression level greater than 37% using the mask-based compression method. Thus, the mask-based compression method operating alone produced some frames with a compression level of less than 37%, but the alternating method provided by embodiments of the present invention limits the minimum compression level to approximately 37%, as shown in Figure 16. For frames with significant black pixel content, the mask-based compression method provides a high level of compression, but for frames with limited black pixel content, the frame-based compression method establishes a lower limit of compression level, for example 37% in this illustrated embodiment. As will be apparent to those skilled in the art, the minimum compression level does not have to be 37%, which is merely illustrative, and other minimum compression levels can be utilized depending on the specific application. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0078]

[0104] Information about the compression method used for each frame can be provided to the endpoint, such as a decoder or display, so that the endpoint can utilize the appropriate decompression method when reconstructing each frame.

[0079]

[0105] Figure 17 is a simplified flowchart illustrating a method for compressing an image frame using an alternating compression algorithm according to one embodiment of the present invention. Method 1700 includes receiving a frame of video data (1710). The method also includes determining the number of lines in the frame that have pixel groups characterized by luminance levels below a threshold (1712).

[0080]

[0106] If the number of lines is greater than or equal to the compression threshold (1714), the frame is compressed using a mask-based compression method (1720). If the number of lines is less than the compression threshold, the frame is compressed using a frame-based compression method (1722). If additional frames exist (1730), the method operates on the next frame of video data by receiving the frame of video data (1710). Otherwise, the method terminates (1740). Thus, embodiments of the present invention alternate between compressing each frame using the respective compression methods, depending on the level of compression that can be achieved by each compression method.

[0081]

[0107] It should be noted that the specific steps shown in Figure 17 provide a particular method for compressing image frames using an alternating compression algorithm according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figure 17 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0082]

[0108] According to some embodiments of the present invention, for each frame, there is an embedded image line control or alternative control mechanism that provides the endpoint display with information on which system should be used to decode the incoming MIPI frame. In addition, a virtual MIPI channel can be used to indicate the compression ratio used by the endpoint display.

[0083]

[0109] Some embodiments of the present invention modify the compression quality based on target tracking, thereby reducing the quality in the foveal region to give a higher compression ratio. This is done for the MIPI interface, thereby reducing the amount of data transmitted to the LCOS / uLED display via MIPI. Thus, the embodiments also result in power saving.

[0084]

[0110] Embodiments of the present invention reduce the amount of stream-based data transmitted via MIPI compression. Furthermore, embodiments modify the compression quality based on target tracking, thereby providing a higher compression ratio to the foveal region where quality is degraded. Moreover, embodiments enable a higher compression ratio for steam-based compression techniques while maintaining quality in the area observed by the user. As a result, embodiments enable a much higher compression ratio while maintaining quality.

[0085]

[0111] For stream-based compression standards such as DSC and VESA display compression (VDC-X), low-latency implementations are utilized. This low-latency response is used so that any previous spatial warp adjustments performed remain applicable.

[0086]

[0112] Figure 18 is a simplified image showing an image frame divided into a high-quality region and a low-quality region according to one embodiment of the present invention. The image 1800 shown in Figure 18 includes a high-quality region 1810 and a low-quality region 1820. As will be discussed in more detail below, the high-quality region 1810 is compressed and decompressed using a first quality setting or compression level, and the low-quality region 1820 or the entire image is compressed and decompressed using a second quality setting or compression level, providing memory savings and other advantages. As an example, a single decoder can be utilized by not compressing the high-quality region 1810 and compressing the low-quality region using a single decoder. Significant savings can be achieved if the high-quality region 1810 is small compared to the entire image. An additional explanation regarding resizing the high-quality region is provided in U.S. Provisional Patent Application No. 63 / 543,876, filed October 12, 2023, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0087]

[0113] DSC

[0114] Conventional DSCs do not offer variable quality compression. Rather, DSCs take 24-bit color coding and compress it to 15 / 12 / 10 / 8 bits. The higher the compression (24→8bpp), the worse the impact on quality. With respect to the quality required for the section the eye is focusing on, embodiments can maintain PSNR quality settings above 60 dB, as discussed above. From the use case analysis shown in Figure 6, the inventors determined that this occurs only at a 37% compression configuration (24→15bpp). However, only the area the eye is currently focusing on actually utilizes that compression setting. The lateral foveal region (e.g., the part of the image further from the line of sight) can have lower quality, for example, a 75% compression level (24→8bpp).

[0088]

[0115] Therefore, in neighborhood-based compression standards like DSC, which lack the concept of tiles, embodiments divide the main screen into a high-quality area and a low-quality area (as shown in Figure 18) or smaller sections (as shown in Figure 20), each with a different compression ratio. The selected compression ratio is a function of the current viewing position. Thus, referring to Figure 18, where the viewing position is positioned inside the high-quality area 1810, the high-quality area 1810 can be compressed at a lower compression level (e.g., 24 → 15 bpp), and the low-quality area 1820 can be compressed at a higher compression level (e.g., 24 → 8 bpp). In some examples, the low-quality area 1820 can be compressed at an even higher compression level (e.g., 24 → 6 bpp). In embodiments where the entire image is compressed using a higher compression level, as will be described more fully herein, the high-quality area 1810 can be superimposed on the entire image when the image is reconstructed.

[0089]

[0116] Figure 19 is a simplified flowchart illustrating a method 1900 for compressing an image using different compression ratios for high-quality and low-quality regions, according to one embodiment of the present invention. Method 1900 includes determining the user's line of sight (1910), generating a foveal map containing a first region of the image and a second region of the image (1912), and compressing the first region using a first compression ratio and the second region using a second compression ratio (1914).

[0090]

[0117] The image may be an image contained within a video stream. Determining the user's gaze position can be achieved by utilizing an eye-tracking system that provides the gaze position as a function of time. The foveal map defines the compression ratio to which portions of the image are compressed, which varies as a function of the position in the image relative to the gaze position, with areas(s) closer to the gaze position being compressed using a lower compression ratio and areas(s) further away from the gaze position being compressed using a higher compression ratio. In the example shown in Figure 18, the foveal map contains two regions, but the present invention is not limited to this particular implementation, and three or more regions can be defined. In some embodiments, the foveal map includes a first region of the image and a second region of the image. Method 1900 may be referred to as N-directional compression (e.g., DSC, VDC-X, or JPEG), where N refers to the number of regions determined for the image. For example, based on the gaze position, high-quality regions, medium-quality regions surrounding the high-quality regions, and low-quality regions can be determined for the image. The technique of Method 1900 can then be used as three-directional compression with different compression ratios for each region.

[0091]

[0118] Referring back to Figure 18, in some examples, the low-quality region 1820 may encompass the entire image, including the portion of the image within the high-quality region 1810 characterized by the gaze position. When decoding a compressed image (for example, for reconstruction for display to a user), it may be desirable to decode sections of the image in parallel. For an image divided into a high-quality region 1810 and a low-quality region 1820, as in Figure 18, the low-quality region 1820 may be considered the entire image. For example, in the case of a 2-kilopixel × 2-kilopixel image (4 megapixels total), the low-quality region 1820 may be the entire 4-megapixel image and may be compressed using a high compression level (e.g., 24 → 8 bpp). The high-quality region 1810 may be determined based on the current gaze position and may be, for example, a 1-kilopixel × 1-kilopixel region (1 megapixel total). The high-quality region 1810 can be compressed using a low compression level (e.g., 24 → 15 bpp). Therefore, two DSC decoders can be used to decode a compressed image. During image reconstruction, the decoded high-quality regions can be overlaid on the decoded low-quality regions.

[0092]

[0119] Figure 20 is a simplified image showing an image frame divided into high-quality and low-quality sections according to one embodiment of the present invention. As will be discussed in more detail below, the divided image frame 2000 shown in Figure 20 can be used to define a foveal map that defines the compression ratio at which different sections of the image are compressed, such that the compression ratio or other compression quality metric varies as a function of the position in the image relative to the gaze position. For example, sections closer to the gaze position can be compressed using a lower compression ratio, and sections further from the gaze position can be compressed using a higher compression ratio.

[0093]

[0120] Referring to Figure 20, four sections 2010, 2012, 2014, and 2016, including the high-quality region 2002 (i.e., the region corresponding to the current gaze position), are compressed at a lower compression level (e.g., 24 → 15 bpp), while the remaining sections, which may be referred to as peripheral sections or low-quality sections, are compressed at a higher compression level (e.g., 24 → 8 bpp). As a result, when the compressed image is reconstructed for display to the user, the high-quality region corresponding to the gaze position is characterized by higher quality than the rest of the image further from the gaze position. Consequently, embodiments of the present invention provide gaze-position-based fovealed images with reduced storage and transmission requirements.

[0094]

[0121] In some embodiments of the example shown in Figure 20, all sections 2010–2046 of the image may be compressed at a high compression ratio (e.g., 24 to 8 bpp). Four sections 2010, 2012, 2014, and 2016 containing high-quality regions may also be compressed at a lower compression ratio (e.g., 24 to 15 bpp). By using a decoder, all sections 2010–2046 compressed at a high compression ratio can be decoded according to a higher compression ratio, and the four sections 2010, 2012, 2014, and 2016 compressed at a lower compression ratio can be decoded according to a lower compression ratio. During image reconstruction, the decoded high-quality sections 2010, 2012, 2014, and 2016 can be superimposed on the decoded low-quality sections 2010–2046. In some embodiments, a foveal map can define sections that coincide with high-quality regions. For example, sections 2010–2016 may include only high-quality regions characterized by gaze position, excluding portions of images within low-quality regions.

[0095]

[0122] Similar to N-direction compression, it may be desirable to decode a compressed image using section-based DSC techniques with multiple DSC decoders. For example, a compressed image can be decoded using four DSC decoders: one decoder used to decode high-quality sections 2010-2016, another used to decode sections 2020-2026, a third used to decode sections 2030-2036, and a fourth used to decode sections 2040-2046, with each decoder using a compression ratio for each group of sections based on its proximity to the viewing position. In some embodiments, a single decoder can be implemented with acceptable latency when decoding a compressed image, depending on the memory capacity (e.g., SRAM) of the system used for decoding.

[0096]

[0123] The image may be an image contained within a video stream. Determining the user's gaze position can be achieved by utilizing an eye-tracking system that provides the gaze position as a function of time. The foveal map defines the compression ratios to which different sections of the image (e.g., sections 2010-2016, 2020-2026, 2030-2036, and 2040-2046) are compressed, which vary as a function of their position in the image relative to the gaze position, with sections closer to the gaze position being compressed using lower compression ratios and sections further away from the gaze position being compressed using higher compression ratios. In the example shown in Figure 20, the foveal map contains 16 sections, but the present invention is not limited to this particular implementation, and more or fewer sections may be defined. The methods described herein may be referred to as section-based compression (e.g., DSC, VDC-X, or JPEG) methods.

[0097]

[0124] While some of the examples above show only two compression levels, embodiments of the present invention are not limited to these specific compression levels, and an additional number of compression levels can be utilized. For example, sections 2010–2014 can be compressed using a 37% compression level (i.e., 24→15bpp), while sections 2020, 2022, 2024, and 2026, which are further from the high-quality region, can be compressed using a 50% compression level (i.e., 24→15bpp), sections 2030, 2032, 2034, and 2036, which are even further from the high-quality region than sections 2020–2026, can be compressed using a 58% compression level (i.e., 24→12bpp), and sections 2040, 2042, 2044, and 2046, which are the furthest from the high-quality region than sections 2010–2016, can be compressed using a 67% compression level (i.e., 24→8bpp). Therefore, the use of two compression levels is merely illustrative. Furthermore, for some sections, the compression level may be 0%, i.e., uncompressed, including sections corresponding to the viewing position and high-quality areas. Thus, a compressed image may have both uncompressed and compressed sections. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0098]

[0125] Furthermore, although Figure 20 shows only 16 sections of uniform area, this is not mandatory, and other numbers of sections with different sizes can be used, with smaller sections adjacent to the high-quality areas, and larger sections, such as those compressed at a higher level, being further away from the high-quality areas. Thus, the number of compression levels, the compression levels, the number of sections, and the size of the sections can vary depending on the specific application. Those skilled in the art will recognize many variations, modifications, and substitutions.

[0099]

[0126] When image compression reduces the frame size, the communication interface, such as the MIPI interface, can be changed to enter a low-power data transmission mode, or even an ultra-low-power sleep mode, thereby saving computing resources and reducing power consumption. At the endpoint, the reconstruction of the compressed image can be performed before it is displayed to the user.

[0100]

[0127] Figure 21 is a simplified flowchart illustrating a method 2100 for compressing an image using different compression ratios for high-quality and low-quality sections, according to one embodiment of the present invention. Method 2100 includes determining the user's line of sight (2110), generating a foveal map containing a first section of the image and a second section of the image (2112), and compressing the first region using a first compression ratio and the second region using a second compression ratio (2114).

[0101]

[0128] It should be recognized that the specific steps shown in Figures 19 and 21 provide a specific method for compressing an image according to one embodiment of the present invention. Other sets of steps may be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Furthermore, the individual steps shown in Figures 19 and 21 may include a plurality of substeps that can be performed in various sequences as appropriate to the individual steps. Additionally, additional steps may be added or removed depending on the specific application. Those skilled in the art will recognize many variations, modifications, and alternatives.

[0102]

[0129] VDC-X

[0130] The VDC-X compression standard (e.g., VDC-M) uses a tile-based method instead of the nearest neighbor method. This compression standard encodes different tiles with different quality settings, but the goal of this conventional compression is to maintain a constant frame size (i.e., bitrate) overall. Therefore, the compression ratio, when selected, varies for each tile to maintain a constant bitrate. Using this compression standard in conjunction with embodiments of the present invention, video images are compressed based on the user's gaze position, not solely on the bitrate. As an example, four sections 2010, 2012, 2014, and 2016, which contain high-quality areas (i.e., areas corresponding to the current gaze position), are compressed with a higher quality setting than the remaining sections, which can be called peripheral sections, and the remaining sections are compressed with a lower quality setting than those used for sections 2010-2016.

[0103]

[0131] In some embodiments of the present invention, since a constant bitrate is not maintained, the size of each frame changes over time, and the transport interface, such as MIPI, enters a low-power mode when not in use.

[0104]

[0132] Similar to the DSC-based method discussed above, in the VDC-X tile-based method, the embodiment encodes the quality of each tile based on the current position of the user's gaze. As shown in Figure 20, using gaze information provided by the AR system's gaze tracking system, the tiles are compressed using the VDC-X standard as a function of the distance of the tile from the gaze position.

[0105]

[0133] Therefore, embodiments of the present invention allow for variations in frame size or per-frame bitrate, and use the current line-of-sight information to select which tiles (VDC-X) or sections (DSC) have higher quality than foveal regions with lower quality settings.

[0106]

[0134] In some embodiments, the N-directional compression or section-based compression described above can implement JPEG as a compression standard rather than DSC or VDC-X. In these embodiments, the compression ratio used for high-quality / low-quality regions and / or high-quality / low-quality sections can instead refer to the quality settings of the JPEG standard.

[0107]

[0135] Figure 22 is a simplified block diagram showing the components of an AR system according to one embodiment of the present invention. The AR system 2200 shown in Figure 22 can be incorporated into an AR device as described herein. Figure 22 provides a schematic diagram of one embodiment of the AR system 2200 that can carry out some or all of the steps of the method provided by various embodiments. It should be noted that Figure 22 is intended only to provide a generalized description of the various components, some or all of which may be appropriately utilized. Thus, Figure 22 broadly illustrates how individual system elements may be implemented in a relatively isolated or relatively more integrated manner.

[0108]

[0136] The AR system 2200 is shown as comprising hardware elements that can be electrically coupled via bus 2205 or otherwise communicate as needed. The hardware elements may include, but are not limited to, one or more processors 2210, including one or more general-purpose processors and / or one or more dedicated processors such as digital signal processing chips, graphics accelerators, and / or the like; one or more input devices 2215, which may include, but are not limited to, a mouse, keyboard, camera, and / or the like; and one or more output devices 2220, which may include, but are not limited to, a display device, printer, and / or the like. Additionally, the AR system 2200 includes an eye-tracking system 2255 that can provide the AR system with the user's gaze position. The foveal image compression techniques considered herein can be implemented using the processor 2210.

[0109]

[0137] The AR system 2200 may further include, but is not limited to, one or more non-temporary storage devices 2225 that may include local and / or network-accessible storage and / or be able to communicate with them, and / or may include, but is not limited to, solid-state storage devices such as random-access memory (RAM) and / or read-only memory (ROM), which may be disk drives, drive arrays, optical storage devices, programmable, flash-updatable and / or similar. Such storage devices may be configured to implement any suitable data store, including, but is not limited to, various file systems, database structures and / or similar.

[0110]

[0138] The AR system 2200 may also include a communication subsystem 2219 which may include, but is not limited to, a chipset that includes a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a Bluetooth® device, an 802.11 device, a WiFi device, a WiMAX device, a cellular communication equipment, and / or similar. The communication subsystem 2219 may include one or more input and / or output communication interfaces to enable exchange of data with a network, such as, to give an example, another computer system, a television, and / or any other device described herein. Depending on the desired functionality and / or other implementation concerns, a portable electronic device or similar device may communicate images and / or other information via the communication subsystem 2219. In other embodiments, a portable electronic device, e.g., a first electronic device, may be incorporated into the AR system 2200, e.g., an electronic device as an input device 2215. In some embodiments, the AR system 2200 further includes a working memory 2260 which may include a RAM or ROM device, as described above.

[0111]

[0139] The AR system 2200 may also include software elements, shown as currently located in working memory 2260, including an operating system 2262, device drivers, executable libraries, and / or computer programs provided by various embodiments, and / or other code such as one or more application programs 2264 that may be designed to implement and / or configure the system as described herein, in accordance with the methods provided by other embodiments. As mere examples, one or more procedures described with respect to the methods considered above may be implemented as code and / or instructions executable by a computer and / or a processor within a computer. In one embodiment, such code and / or instructions may be used to configure and / or adapt a general-purpose computer or other device to perform one or more operations in accordance with the described methods.

[0112]

[0140] These instructions and / or sets of code can be stored in a non-temporary computer-readable storage medium such as the storage device(s) 2225 described above. In some cases, the storage medium may be incorporated into a computer system such as the AR system 2200. In other embodiments, the storage medium may be separate from the computer system, and may be a removable medium such as a compact disk, and / or provided in an installation package, so that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the stored instructions / code. These instructions may take the form of executable code that can be executed by the AR system 2200, and / or in the form of source and / or installable code, which takes the form of executable code when compiled and / or installed on the AR system 2200 using, for example, one of various commonly available compilers, installers, compression / decompression utilities, etc.

[0113]

[0141] It will be apparent to those skilled in the art that substantial modifications can be made according to specific requirements. For example, customized hardware may be used, and / or certain elements may be implemented in hardware, software including portable software such as applets, or both. Furthermore, connections to other computing devices, such as network input / output devices, may be used.

[0114]

[0142] As described above, in one embodiment, several embodiments can be implemented using a computer system such as AR system 2200 to carry out methods according to various embodiments of the Art. According to a set of embodiments, some or all of the steps of such methods are carried out by AR system 2200 in response to a processor 2210 that executes one or more sequences of one or more instructions that may be incorporated into an operating system 2262 and / or other code such as an application program 2264, which are contained in working memory 2260. Such instructions may be read into working memory 2260 from another computer-readable medium, such as one or more of storage devices 2225. As just one example, the execution of a sequence of instructions contained in working memory 2260 can cause the processor 2210 to carry out one or more steps of the methods described herein. Additionally or alternatively, some of the methods described herein may be carried out via dedicated hardware.

[0115]

[0143] As used herein, the terms machine-readable medium and computer-readable medium refer to any medium involved in providing data that causes a machine to operate in a particular way. In embodiments implemented using the AR system 2200, various computer-readable media may be involved in providing instructions / code to the processor(s) 2210 for execution, and / or may be used to store and / or carry such instructions / code. In many implementations, the computer-readable medium is a physical and / or tangible storage medium. Such media may take the form of non-volatile or volatile media. Non-volatile media include, for example, optical and / or magnetic disks such as the storage device(s) 2225. Volatile media include, but are not limited to, dynamic memory such as the working memory 2260.

[0116]

[0144] Common forms of physical and / or tangible computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH-EPROMs, any other memory chips or cartridges, or any other media from which a computer can read instructions and / or code.

[0117]

[0145] Various forms of computer-readable media may be involved in transporting one or more sequences of one or more instructions to the processor(s) 2210 for execution. As just one example, the instructions may first be transported onto a magnetic disk and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions as signals over a transmission medium to be received and / or executed by the AR system 2200.

[0118]

[0146] The communication subsystem 2219 and / or its components generally receive signals, and the bus 2205 may then transport the signals and / or the data, instructions, etc. carried by the signals to the working memory 2260, from which the processor(s) 2210 retrieves and executes the instructions. Instructions received by the working memory 2260 may optionally be stored in a non-temporary storage device 2225 either before or after execution by the processor(s) 2210.

[0119]

[0147] Various embodiments of the present disclosure are provided below. When used herein, any reference to a set of embodiments should be understood as a disjunctive reference to each of those embodiments (for example, “Embodiments 1-4” should be understood as “Embodiments 1, 2, 3, or 4”).

[0120]

[0148] Example 1 is a method for compressing an image, comprising: determining the user's gaze position; generating a foveal map based on the gaze position, wherein the foveal map includes a first region of the image and a second region of the image; and compressing the first region of the image using a first quality setting and compressing the second region of the image using a second quality setting.

[0121]

[0149] Example 2 is the same as in Example 1, wherein determining the gaze position involves using an eye-tracking camera of an augmented reality device.

[0122]

[0150] Example 3 is the method described in Examples 1-2, wherein the foveal map includes the central region and the peripheral region.

[0123]

[0151] Example 4 is the method according to Examples 1 to 3, wherein the image includes virtual content generated by an augmented reality device.

[0124]

[0152] Example 5 is the same as the method described in Examples 1 to 4, wherein the image is included in a virtual content video stream.

[0125]

[0153] Example 6 is the method according to Examples 1 to 5, wherein compressing a first region of an image using a first quality setting is equivalent to compressing all blocks within the first region using a first quality setting.

[0126]

[0154] Example 7 is the method described in Examples 1 to 6, wherein the first quality setting is higher than the second quality setting.

[0127]

[0155] Example 8 is the method described in Examples 1 to 7, wherein the first quality setting is 100%.

[0128]

[0156] Example 9 is the method according to Examples 1 to 8, further comprising post-processing image content in at least one of the first or second region.

[0129]

[0157] Example 10 is a method according to Examples 1 to 9, further comprising compressing to generate a compressed image and decoding the compressed image using a foveal map.

[0130]

[0158] Example 11 is the method according to Examples 1 to 10, wherein a first region of the image comprises a plurality of first blocks, a second region of the image comprises a plurality of second blocks, compressing the first region of the image comprises compressing each of the plurality of first blocks using a first quality setting, and compressing the second region of the image comprises compressing each of the plurality of second blocks using a second quality setting.

[0131]

[0159] Example 12 is the method according to Examples 1 to 11, further comprising decompressing a first region of an image using a first quality setting, decompressing a second region of an image using a second quality setting, and displaying the image to the user.

[0132]

[0160] Example 13 is the method according to Examples 1 to 12, wherein the second region of the image includes the first region of the image.

[0133]

[0161] Example 14 is a method of Examples 1 to 13, further comprising: compressing to generate a compressed image; decoding the compressed image using a foveal map to generate a decoded first region and a decoded second region; and reconstructing the image by overlaying the decoded first region on top of the decoded second region.

[0134]

[0162] Embodiment 15 is an augmented reality (AR) system comprising: a wearable device, which includes a frame, a projector coupled to the frame, a display optically coupled to the projector, and an eye-tracking system; a memory; and a processor, which is configured to receive eye-tracking position from the eye-tracking system, generate an image, generate a foveal map based on the eye-tracking position, wherein the foveal map includes a first region of the image and a second region of the image, compress the first region of the image using a first quality setting, and compress the second region of the image using a second quality setting.

[0135]

[0163] Example 16 is the AR system described in Example 15, wherein the projector includes one projector from a set of projectors, the display includes one display from a set of displays, and the eye-tracking system includes a set of eye-tracking devices.

[0136]

[0164] Example 17 is an AR system described in Examples 15-16, wherein determining the gaze position involves the use of an eye-tracking camera of an augmented reality device.

[0137]

[0165] Example 18 is an AR system described in Examples 15-17, wherein the foveal map includes a central region and a peripheral region.

[0138]

[0166] Example 19 is an AR system according to Examples 15-18, wherein the image includes virtual content generated by an augmented reality device.

[0139]

[0167] Example 20 is an AR system described in Examples 15-19, wherein the image is included in a virtual content video stream.

[0140]

[0168] Example 21 is an AR system according to Examples 15-20, wherein compressing a first region of an image using a first quality setting includes compressing all blocks within the first region using a first quality setting.

[0141]

[0169] Example 22 is an AR system described in Examples 15-21, in which the first quality setting is higher than the second quality setting.

[0142]

[0170] Example 23 is an AR system described in Examples 15-22, wherein the first quality setting is 100%.

[0143]

[0171] Example 24 is an AR system according to Examples 15-23, wherein the processor is further configured to post-process image content in at least one of the first or second regions.

[0144]

[0172] Example 25 is an AR system according to Examples 15-24, further configured to compress an image, generate a compressed image, and have a processor decode the compressed image using a foveal map.

[0145]

[0173] Example 26 is an AR system according to Examples 15 to 25, wherein a first region of the image comprises a plurality of first blocks, a second region of the image comprises a plurality of second blocks, compressing the first region of the image comprises compressing each of the plurality of first blocks using a first quality setting, and compressing the second region of the image comprises compressing each of the plurality of second blocks using a second quality setting.

[0146]

[0174] Example 27 is an AR system according to Examples 15-26, wherein the processor is further configured to decompress a first region of the image using a first quality setting, decompress a second region of the image using a second quality setting, and display the image to the user.

[0147]

[0175] Example 28 is an AR system according to Examples 15-27, wherein the second region of the image includes the first region of the image.

[0148]

[0176] Example 29 is an AR system according to Examples 15-28, further configured such that compression generates a compressed image, a processor decodes the compressed image using a foveal map to generate a decoded first region and a decoded second region, and the image is reconstructed by overlaying the decoded first region on top of the decoded second region.

[0149]

[0177] Embodiment 30 is a non-temporary computer-readable medium comprising program code executable by a processor of a user-wearable device, wherein the program code is executable by the processor to determine the user's gaze position, generate a foveal map based on the gaze position, the foveal map comprising a first region of an image and a second region of an image, compress the first region of the image using a first quality setting and compress the second region of the image using a second quality setting.

[0150]

[0178] In the aforementioned specification, this disclosure is described with reference to specific embodiments thereof. However, it will be apparent that various modifications and changes can be made without departing from the broader spirit and scope of this disclosure. Accordingly, this specification and the drawings should be considered illustrative rather than restrictive.

[0151]

[0179] In fact, it will be recognized that each of the systems and methods of this disclosure has several innovative aspects, and that not one alone alone can assume or be required for the desired attributes disclosed herein. The various features and processes described above may be used independently of each other or combined in various ways. All possible combinations and partial combinations are intended to fall within the scope of this disclosure.

[0152]

[0180] Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any preferred partial combination in multiple embodiments. Furthermore, features described above as acting in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be removed from the combination, and the claimed combination may cover a partial combination or a variation of a partial combination. Not a single feature or group of features is required or essential in every embodiment.

[0153]

[0181] In particular, conditional statements used herein, such as “can,” “could,” “might,” “may,” and “for example,” should generally be understood, unless otherwise specified or understood in the context in which they are used, to convey that a particular embodiment includes certain features, elements, and / or steps, but other embodiments do not. Therefore, such conditional statements are not generally intended to mean that features, elements, and / or steps are required in some way in one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are included in any particular embodiment, or should be implemented in a particular embodiment, with or without input or instruction from the author. Terms such as “comprising,” “including,” and “having” are synonymous and are used comprehensively and openly, without prejudice to additional elements, features, actions, behaviors, etc. Furthermore, the term "or" is used in its inclusive sense (rather than its exclusive sense), so for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In addition, as used in this application and the attached claims, the articles "a," "an," and "the" should be interpreted as meaning "one or more" or "at least one" unless otherwise specified. Similarly, while actions may be depicted in a particular order in the drawings, it should be recognized that, in order to achieve the desired result, such actions do not need to be performed in a particular order or sequential order shown, or that all shown actions will be performed. Moreover, the drawings may schematically illustrate another exemplary process in the form of a flowchart. However, other actions not depicted may be incorporated into the schematicly shown exemplary methods and processes. For example, one or more additional actions may be performed before, after, simultaneously with, or in between any of the illustrated actions.Additionally, the operations may be configured or ordered in other embodiments. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other embodiments are within the scope of the following claims. In some cases, the operations described in the claims may be performed in a different order and still achieve the desired results.

[0154]

[0182] Accordingly, the claims are not intended to be limited to the embodiments shown herein, but should be given the broadest scope consistent with this disclosure, the principles and novel features disclosed herein. Accordingly, it should be understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or changes in light thereof should be suggested to those skilled in the art and should be included within the spirit and scope of this application and the appended claims.

Claims

1. A method of compressing images, Determining the user's line of sight, The method involves generating a foveal map based on the aforementioned line of sight position, wherein the foveal map includes a first region of the image and a second region of the image. Compressing the first region of the image using a first quality setting, and compressing the second region of the image using a second quality setting, Methods that include...

2. The method according to claim 1, wherein determining the gaze position includes using an eye-tracking camera of an augmented reality device.

3. The method according to claim 1, wherein the foveal map includes a central region and a peripheral region.

4. The method according to claim 1, wherein the image includes virtual content generated by an augmented reality device.

5. The method according to claim 4, wherein the image is included in a virtual content video stream.

6. The method according to claim 1, wherein compressing the first region of the image using the first quality setting includes compressing all blocks within the first region using the first quality setting.

7. The method according to claim 1, wherein the first quality setting is higher than the second quality setting.

8. The method according to claim 7, wherein the first quality setting is 100%.

9. The method according to claim 1, further comprising post-processing image content in at least one of the first region or the second region.

10. The method according to claim 1, wherein a compressed image is generated by the compression, and the method further comprises decoding the compressed image using the foveal map.

11. The first region of the aforementioned image includes a plurality of first blocks, The second region of the aforementioned image includes a plurality of second blocks, Compressing the first region of the image includes compressing each of the plurality of first blocks using the first quality setting, The method according to claim 1, wherein compressing the second region of the image comprises compressing each of the plurality of second blocks using the second quality setting.

12. Decompressing the first region of the image using the first quality setting, Decompressing the second region of the image using the second quality setting, Displaying the aforementioned image to the user, The method according to claim 1, further comprising:

13. The method according to claim 1, wherein the second region of the image includes the first region of the image.

14. The above compression generates a compressed image, and the above method, The compressed image is decoded using the foveal map to generate a decoded first region and a decoded second region. The image is reconstructed by superimposing the decoded first region onto the decoded second region. The method according to claim 13, further comprising:

15. It is an augmented reality (AR) system, It is a wearable device, Frame and, A projector coupled to the aforementioned frame, A display optically coupled to the aforementioned projector, Eye-tracking system and, Wearable devices, Memory and It is a processor, Receiving the gaze position from the aforementioned gaze tracking system, To generate an image, The method involves generating a foveal map based on the aforementioned line of sight position, wherein the foveal map includes a first region of the image and a second region of the image. Compressing the first region of the image using a first quality setting, and compressing the second region of the image using a second quality setting, A processor configured to perform the following actions: An augmented reality (AR) system equipped with [the following features].

16. The AR system according to claim 15, wherein the projector includes one projector from a set of projectors, the display includes one display from a set of displays, and the eye-tracking system includes a set of eye-tracking devices.

17. The AR system according to claim 16, wherein determining the gaze position includes using an eye-tracking camera of an augmented reality device.

18. The AR system according to claim 16, wherein the foveal map includes a central region and a peripheral region.

19. The AR system according to claim 16, wherein the image includes virtual content generated by an augmented reality device.

20. The AR system according to claim 19, wherein the aforementioned image is included in a virtual content video stream.

21. The AR system according to claim 16, wherein compressing the first region of the image using the first quality setting includes compressing all blocks within the first region using the first quality setting.

22. The AR system according to claim 16, wherein the first quality setting is higher than the second quality setting.

23. The AR system according to claim 22, wherein the first quality setting is 100%.

24. The AR system according to claim 16, wherein the processor is further configured to post-process image content in at least one of the first region or the second region.

25. The AR system according to claim 16, wherein compression generates a compressed image, and the processor is further configured to decode the compressed image using the foveal map.

26. The first region of the aforementioned image includes a plurality of first blocks, The second region of the aforementioned image includes a plurality of second blocks, Compressing the first region of the image includes compressing each of the plurality of first blocks using the first quality setting, The AR system according to claim 16, wherein compressing the second region of the image comprises compressing each of the plurality of second blocks using the second quality setting.

27. The aforementioned processor, Using the first quality setting, decompress the first region of the image, Using the second quality setting, decompress the second region of the image, Display the aforementioned image to the user. The AR system according to claim 16, further configured as follows.

28. The AR system according to claim 16, wherein the second region of the image includes the first region of the image.

29. Compression generates a compressed image, and the processor, The compressed image is decoded using the foveal map to generate a decoded first region and a decoded second region. The image is reconstructed by superimposing the decoded first region onto the decoded second region. The AR system according to claim 15, further configured as follows.

30. A non-temporary computer-readable medium comprising program code executable by the processor of a device that can be worn by a user, wherein the program code is executed by the processor, Determining the user's line of sight, The method involves generating a foveal map based on the aforementioned line of sight position, wherein the foveal map includes a first region of the image and a second region of the image. Compressing the first region of the image using a first quality setting, and compressing the second region of the image using a second quality setting, A non-temporary, computer-readable medium that is executable to perform the following actions.