Method and system for performing gaze point image compression based on eye gaze
By combining foveated image compression technology with an eye-tracking system, the problems of resource waste and poor user experience in augmented reality systems have been solved, achieving efficient image quality management and improved user experience.
Patent Information
- Application Number
- CN202480020041.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-20
- Filing Date
- 2024-03-19
- Publication Date
- 2025-11-07
AI Technical Summary
Existing augmented reality systems have failed to effectively integrate human visual characteristics into their display systems, resulting in wasted resources and poor user experience.
A gaze-point image compression method is adopted. The user's gaze position is determined by an eye tracking system. The gaze-point image compression technology is used to reduce the image quality of the non-gaze area and improve the image quality of the gaze area.
It reduces the need for storage and processing resources while improving the user experience and maintaining high-quality display in the gaze area.
Smart Images

Figure CN120917751A_ABST
Abstract
Description
Cross Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 453,376, filed March 20, 2023, entitled “Method and System for Performing Foveated Image Compression Based on Eye Gaze,” the disclosure of which is hereby incorporated by reference in its entirety for all purposes. BACKGROUND Modern computing and display technologies have facilitated the development of systems for so called virtual reality or augmented reality experiences, in which digitally reproduced images or portions of images are presented to a viewer in a manner wherein they seem to be, or can be perceived as, real. A virtual reality, or VR, scenario typically involves presentation of digital or virtual image information only, e.g., a “head-mounted display” or headset — a helmet or frame- like device which a viewer wears on their head, eye(s), and / or ear(s), to present digital or virtual image information. An augmented reality, or AR, scenario typically involves presentation of both digital or virtual image information and actual real-world image information, e.g., via a transparent display device having a digital or virtual image projected into the viewer’s field of view. REFERENCE Figure 1 An augmented reality scenario can also involve multiple users with multiple associated AR devices. A user with an AR device may, for example, interact with other users and / or with elements of a real-world scene, represented by real-world image information, e.g., video, audio, etc., and / or with elements of a digital or virtual scene, represented by digital or virtual image information, e.g., video, audio, etc. Despite advances in these display technologies, there remains a need for improvements in methods and systems related to augmented reality systems, particularly display systems. SUMMARY The present invention relates generally to methods and systems related to projection display systems including wearable displays. More specifically, embodiments of the present invention provide methods and systems that combine the concept of foveated images (i.e., reducing video quality at a region where the human eye is not focused) with the concept of compression. The present invention is applicable to a variety of applications in computer vision and image display systems, and light field projection systems, including stereoscopic systems, systems that deliver beams of light to a user’s retina, etc. Compared to conventional techniques, the present invention achieves numerous benefits. For example, embodiments of the present invention provide methods and systems capable of compressing portions of an image or video stream corresponding to the user's eye gaze location using higher quality settings than portions of the image or video stream farther from the location corresponding to the user's eye gaze. Therefore, memory and processing resources can be saved, while reducing or minimizing the impact on the user experience. These and other embodiments of the invention, along with their numerous advantages and features, will be described in more detail below and in conjunction with the accompanying drawings. Attached Figure Description Figure 1 This shows the user's view of augmented reality (AR) through an AR device. Figure 2A A cross-sectional side view of an example of a set of stacked waveguides is shown, wherein each waveguide includes an in-line coupling optical element. Figure 2B It shows Figure 2A A perspective view of one or more stacked waveguides in the example. Figure 2C It shows Figure 2A and 2B A top-down plan view of one or more stacked waveguides in the example. Figure 3 This is a simplified diagram of an eyepiece waveguide with a combined pupil expander according to an embodiment of the present invention. Figure 4 An example of a wearable display system according to an embodiment of the present invention is shown. Figure 5 A perspective view of a wearable device according to an embodiment of the present invention is shown. Figure 6 This is a diagram illustrating the run-length encoding of a quantized DCT block according to an embodiment of the present invention. Figure 7 This is a diagram showing the structure of the JPEG header. Figure 8 It is a line graph showing the image compressed using a single quality setting. Figure 9 This is a line drawing illustrating a gaze point image with three gaze point regions according to an embodiment of the present invention. Figure 10 This is a line drawing illustrating a gaze point image undergoing post-processing in the gaze point region according to another embodiment of the present invention. Figure 11 This is a gaze-point 3D generated image with three gaze-point regions according to another embodiment of the present invention. Figure 12 This is a line drawing illustrating an image that can be used in conjunction with multiple gaze mapping maps according to an embodiment of the present invention. Figure 13 is a simplified flowchart illustrating a method of compressing an image according to an embodiment of the application. Figure 14 is a simplified schematic diagram illustrating a gaze-based image gaze point system according to an embodiment of the application. Figure 15 shows the compression level as a function of time for the sparse compression system implementation and the DSC-SPARSE system implementation according to an embodiment of the application, expressed in terms of consecutive frames versus frequency. Figure 16 shows a histogram of the frame count versus compression for the sparse compression system implementation and the DSC-SPARSE system implementation according to an embodiment of the application. Figure 17 is a simplified flowchart illustrating a method of compressing an image frame using an alternate compression algorithm according to an embodiment of the application. Figure 18 is a simplified image illustrating an image frame divided into high quality regions and low quality regions according to an embodiment of the application. Figure 19 is a simplified flowchart illustrating a method of compressing an image using different compression ratios for high quality regions and low quality regions according to an embodiment of the application. Figure 20 is a simplified image illustrating an image frame divided into high quality tiles and low quality tiles according to an embodiment of the application. Figure 21 is a simplified flowchart illustrating a method of compressing an image using different compression ratios for high quality tiles and low quality tiles according to an embodiment of the application. Figure 22 is a simplified block diagram illustrating components of an AR system according to an embodiment of the application. DETAILED DESCRIPTION Reference will now be made to the drawings wherein like numerals refer to like components throughout. The drawings are schematic representations Referring now to the drawings Figure 2A In some embodiments, light incident on a waveguide can need to be redirected to in-couple the light into the waveguide. The light can be redirected and in-coupled into its corresponding waveguide using in-coupling optical elements. Although referred to as “in-coupling optical elements” in this specification, the in-coupling optical elements need not be optical elements, and can be non-optical elements. Figure 2AA cross-sectional side view showing an example of a set of stacked waveguides 200, each containing an in-coupling optical element. Each waveguide can be configured to output light of one or more different wavelengths, or one or more different wavelength ranges. Light from projectors is injected into the set of stacked waveguides 200 and out-coupled to a user, as described in more detail below. The illustrated set of stacked waveguides 200 includes waveguides 202, 204, and 206. Each waveguide includes an associated in-coupling optical element (which can also be referred to as a light input region on the waveguide), where, for example, in-coupling optical element 203 is disposed on a major surface (e.g., an upper major surface) of waveguide 202, in-coupling optical element 205 is disposed on a major surface (e.g., an upper major surface) of waveguide 204, and in-coupling optical element 207 is disposed on a major surface (e.g., an upper major surface) of waveguide 206. In some embodiments, one or more of the in-coupling optical elements 203, 205, 207 can be disposed on a bottom major surface of the respective waveguide 202, 204, 206 (particularly where one or more of the in-coupling optical elements is a reflective, deflection-type in-coupling optical element). As shown, the in-coupling optical elements 203, 205, 207 can be disposed on an upper major surface of their respective waveguide 202, 204, 206 (or the top of the waveguide below), particularly where these in-coupling optical elements are transmissive, deflection-type optical elements. In some embodiments, the in-coupling optical elements 203, 205, 207 can be disposed within the body of the respective waveguide 202, 204, 206. In some embodiments, the in-coupling optical elements 203, 205, 207 are wavelength selective, as described herein, such that they selectively redirect light of one or more wavelengths while transmitting light of other wavelengths. While shown on one side or corner of their respective waveguide 202, 204, 206, it will be appreciated that, in some embodiments, the in-coupling optical elements 203, 205, 207 can be disposed in other regions of their respective waveguide 202, 204, 206. As shown, the in-coupling optical elements 203, 205, 207 can be laterally offset from one another. In some embodiments, each in-coupling optical element can be offset such that it receives light without passing through another in-coupling optical element. For example, each in-coupling optical element 203, 205, 207 can be configured to receive light from a different projector and can be separated (e.g., laterally spaced apart) from the other in-coupling optical elements 203, 205, 207 such that it does not substantially receive light from the other in-coupling optical elements 203, 205, 207. Each waveguide also includes an associated light distributing element, e.g., light distributing element 210 is disposed on a major surface (e.g., a top major surface) of waveguide 202, light distributing element 212 is disposed on a major surface (e.g., a top major surface) of waveguide 204, and light distributing element 214 is disposed on a major surface (e.g., a top major surface) of waveguide 206. In some other embodiments, light distributing elements 210, 212, 214 can be disposed on the bottom major surfaces of the associated waveguides 202, 204, 206, respectively. In some other embodiments, light distributing elements 210, 212, 214 can be disposed on both the top and bottom major surfaces of the associated waveguides 202, 204, 206, respectively; or light distributing elements 210, 212, 214 can be disposed on different ones of the top and bottom major surfaces of the different associated waveguides 202, 204, 206, respectively. Waveguides 202, 204, 206 can be spaced apart and separated by, e.g., layers of gaseous, liquid and / or solid materials. For example, as shown, layer 208 can separate waveguides 202 and 204; layer 209 can separate waveguides 204 and 206. In some embodiments, layers 208 and 209 are formed of a low index of refraction material (i.e., a material having a lower index of refraction than the material forming the immediately adjacent one of waveguides 202, 204, 206). Preferably, the index of refraction of the material forming layers 208 and 209 is 0.05 or more, or 0.10 or less, of the index of refraction of the material forming waveguides 202, 204, 206. Advantageously, lower index layers 208 and 209 can act as cladding layers to facilitate total internal reflection (TIR) of light as it passes through waveguides 202, 204, 206 (e.g., TIR between the top and bottom major surfaces of each waveguide). In some embodiments, layers 208 and 209 are formed of air. Although not shown, it will be appreciated that the top and bottom of the illustrated set of waveguides 200 can include immediate cladding layers. Preferably, for ease of manufacturing and other considerations, the materials forming waveguides 202, 204, 206 are similar or the same, and the materials forming layers 208, 209 are similar or the same. In some embodiments, the materials forming waveguides 202, 204, 206 can differ between one or more waveguides, and / or the materials forming layers 208, 209 can differ, while still maintaining the various index of refraction relationships described above. With continued reference to Figure 2A , light rays 218, 219, 220 are incident on the set of waveguides 200. It will be appreciated that light rays 218, 219, 220 can be injected into waveguides 202, 204, 206 by one or more projectors (not shown). In some embodiments, the light rays 218, 219, and 220 have different properties, such as different wavelengths or different wavelength ranges, which correspond to different colors. The in-line coupling optics 203, 205, and 207 each deflect the incident light, causing it to propagate via TIR through a corresponding waveguide among waveguides 202, 204, and 206. In some embodiments, the in-line coupling optics 203, 205, and 207 each selectively deflect one or more specific wavelengths of light while transmitting other wavelengths to the underlying waveguide and the associated in-line coupling optics.
[0001] For example, the input coupling optical element 203 can be configured to deflect light 218 having a first wavelength or wavelength range, while transmitting light 219 and 220 having different second and third wavelengths or wavelength ranges, respectively. The transmitted light 219 is incident on and deflected by the input coupling optical element 205, which is configured to deflect light of the second wavelength or wavelength range. Light 220 is deflected by the input coupling optical element 207, which is configured to selectively deflect light of the third wavelength or wavelength range. Continue to refer to Figure 2A The deflected light rays 218, 219, and 220 are deflected, causing them to propagate through the corresponding waveguides 202, 204, and 206. That is, the input coupling optical elements 203, 205, and 207 of each waveguide deflect the light into the corresponding waveguide 202, 204, and 206, thereby coupling the light into the corresponding waveguide. The light rays 218, 219, and 220 are deflected at a certain angle, allowing the light to propagate through the corresponding waveguides 202, 204, and 206 via TIR propagation. The light rays 218, 219, and 220 propagate through the corresponding waveguides 202, 204, and 206 via TIR propagation until they strike the corresponding light distribution elements 210, 212, and 214 of the waveguide, where they are output coupled to provide the output coupled light ray 216. Now for reference Figure 2B , showed Figure 2A A perspective view of an example of stacked waveguides. As described above, input coupling rays 218, 219, and 220 are deflected by input coupling optics 203, 205, and 207, respectively, and then propagate within waveguides 202, 204, and 206 via total internal reflection (TIR). Rays 218, 219, and 220 are then incident on light distribution elements 210, 212, and 214, respectively. Light distribution elements 210, 212, and 214 deflect rays 218, 219, and 220, causing them to propagate toward output coupling optics 222, 224, and 226, respectively. In some embodiments, the light distribution elements 210, 212, 214 are orthogonal pupil expanders (OPEs). In some embodiments, the OPEs deflect or distribute light to the out-coupling optical elements 222, 224, 226 and, in some embodiments, can also increase the spot size or beam of the light as it propagates to the out-coupling optical elements. In some embodiments, the light distribution elements 210, 212, 214 can be omitted and the in-coupling optical elements 203, 205, 207 can be configured to deflect light directly to the out-coupling optical elements 222, 224, 226. For example, with reference to Figure 2A , the light distribution elements 210, 212, 214 can be replaced by the out-coupling optical elements 222, 224, 226, respectively. In some embodiments, the out-coupling optical elements 222, 224, 226 are exit pupils (EPs) or exit pupil expanders (EPEs) for directing light to the user’s eye. It can be appreciated that OPEs can be configured to increase the size of the eye box in at least one axis, while EPEs can be configured to increase the eye box in an axis that is crossed (e.g., orthogonal) to the axis of the OPEs. For example, each OPE can be configured to redirect a portion of the light that hits the OPE to the EPE of the same waveguide, while allowing the remainder of the light to continue to propagate down the waveguide. Upon hitting the OPE again, another portion of the remaining light is redirected to the EPE, and the remainder of this portion of light continues to propagate down the waveguide, and so on. Similarly, when hitting the EPE again, a portion of the impinging light is directed out of the waveguide, towards the user, while the remainder of this light continues to propagate down the waveguide until it hits the EPE again, at which point another portion of the impinging light is directed out of the waveguide, and so on. Thus, each time the OPE or EPE redirects a portion of the light, a single in-coupled light beam can be “copied”, forming a field of cloned light beams. In some embodiments, the OPEs and / or EPEs can be configured to modify the size of the light beams. In some embodiments, the functions of the light distribution elements 210, 212, and 214 and the out-coupling optical elements 222, 224, 226 are combined in a combined pupil expander, as discussed with respect to FIG. 2E. Thus, with reference to Figure 2A and 2BIn some embodiments, for each constituent color, a set of waveguides 200 includes waveguides 202, 204, 206; in-coupling optical elements 203, 205, 207; light distributing elements (e.g., OPEs) 210, 212, 214; and out-coupling optical elements (e.g., EPs) 222, 224, 226. The waveguides 202, 204, 206 can be stacked with air gaps / cladding between each waveguide. The in-coupling optical elements 203, 205, 207 redirect or deflect incoming light (different in-coupling optical elements receive different wavelengths of light) into their waveguides. The light then propagates at an angle that will result in TIR within the respective waveguide 202, 204, 206. In the example shown, light ray 218 (e.g., blue light) is deflected by the first in-coupling optical element 203 and then continues to bounce down the waveguide, interacting with the light distributing element (e.g., OPE) 210 in the manner previously described, and then with the out-coupling optical element (e.g., EP) 222. Light rays 219 and 220 (e.g., green and red light, respectively) will pass through the waveguide 202, with light ray 219 hitting and being deflected by the in-coupling optical element 205. Light ray 219 then bounces down the waveguide 204 via TIR, continuing to its light distributing element (e.g., OPE) 212, and then to the out-coupling optical element (e.g., EP) 224. Finally, light ray 220 (e.g., red light) passes through the waveguide 206 to hit the light in-coupling optical element 207 of the waveguide 206. The light in-coupling optical element 207 deflects the light ray 220 so that it propagates via TIR to the light distributing element (e.g., OPE) 214, and then via TIR to the out-coupling optical element (e.g., EP) 226. The out-coupling optical element 226 then finally out-couples the light ray 220 to the viewer, who also receives out-coupled light from the other waveguides 202, 204. Figure 2C A top-down plan view of an example of stacked waveguides in Figure 2A and 2B As shown, the waveguides 202, 204, 206, and the associated light distributing elements 210, 212, 214 and associated out-coupling optical elements 222, 224, 226 of each waveguide can be vertically aligned. However, as described herein, the in-coupling optical elements 203, 205, 207 are not vertically aligned; rather, the in-coupling optical elements are preferably not overlapping (e.g., laterally spaced apart as shown in the top-down or plan view). As described further herein, this non-overlapping spatial arrangement facilitates the one-to-one injection of light from different sources into different waveguides, allowing a particular light source to be uniquely coupled to a particular waveguide. In some embodiments, an arrangement including non-overlapping spatially separated in-coupling optical elements can be referred to as a shifted pupil system, and the in-coupling optical elements within these arrangements can correspond to sub-pupils.
[0002] A simplified diagram of an eyepiece waveguide with a combined pupil expander according to an embodiment of the present application. In Figure 3 In the example shown, eyepiece 310 employs a single-sided configuration of the combined OPE / EPE region. Referring to Figure 3 , eyepiece 310 includes a substrate 320 with an incoupling optical element 322 and a combined OPE / EPE region 324 (also referred to as a combined pupil expander (CPE)) disposed therein. An incident light ray 330 is incoupled via incoupling optical element 320 and outcoupled as an output light ray 332 via combined OPE / EPE region 324. Combined OPE / EPE region 324 includes gratings corresponding to the OPE and EPE that are spatially overlapping in the x- and y-directions. In some embodiments, the gratings corresponding to the OPE and EPE are located on the same side of substrate 320 such that the OPE grating is superimposed on the EPE grating, or vice versa (or both). In other embodiments, the OPE grating is located on an opposite side of substrate 320 from the EPE grating such that the gratings are spatially overlapping in the x- and y-directions but separated from one another in the z-direction (i.e., located in different planes). Thus, combined OPE / EPE region 324 can be implemented in either a single-sided configuration or a double-sided configuration. Figure 4 An example of a wearable display system 430 is shown, in which various waveguides and related systems disclosed herein can be integrated. Referring to Figure 4The display system 430 includes a display 432 and various mechanical and electronic modules and systems to support functioning of the display 432. The display 432 can be coupled to a frame 434, which can be worn by a display system user 440 (also referred to as a viewer), and which is configured to position the display 432 in front of the eyes of the user 440. In some embodiments, the display 432 can be considered eyeglasses. In some embodiments, a speaker 436 is coupled to the frame 434 and is configured to be positioned near an ear canal of the user 440 (in some embodiments, another, not shown, speaker can optionally be positioned near the other ear canal of the user to provide stereo / shapeable sound control). The display system 430 can also incorporate one or more microphones or other devices for detecting sound. In some embodiments, the microphones are configured to allow the user to provide input or commands to the system 430 (e.g., to select voice menu commands, natural language queries, etc.), and / or can allow for audio communication with other people (e.g., with other users of similar display systems). The microphones can also be configured as peripheral sensors for gathering audio data (e.g., sound from the user and / or the environment). In some embodiments, the display system 430 can also incorporate one or more outwardly directed environmental sensors for detecting objects, stimuli, people, animals, locations, or other aspects of the world around the user. For example, the environmental sensors can include one or more cameras, which can be located, for example, in outwardly facing positions so as to capture images similar to at least a portion of the normal field of view of the user 440. In some embodiments, the display system can also include peripheral sensors that can be separate from the frame 434 and attached to the body of the user 440 (e.g., on the head, torso, limbs, etc. of the user 440). In some embodiments, the peripheral sensors can be configured to take data characterizing the physiological state of the user 440. For example, the sensors can be electrodes. The display 432 is operatively coupled by a communication link, such as a wired wire or wireless connection, to a local data processing module, which can be mounted in a variety of configurations, such as fixedly attached to the frame 434, fixedly attached to a helmet or hat worn by the user, embedded in an earpiece, or otherwise removably attached to the user 440 (e.g., in a backpack-style configuration, a band-coupled configuration). Similarly, the sensors can be operatively coupled by a communication link, such as a wired wire or wireless connection, to a local processor and data module. The local processing and data module can include a hardware processor as well as digital memory, such as nonvolatile memory (e.g., flash memory or solid state disk), both of which can be used to assist in the processing, caching, storage, and / or retrieval of data. The local processor and data module can include one or more central processing units (CPU), graphics processing units (GPU), dedicated processing hardware, and the like. Data can include a) data captured from sensors, such as image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, radio devices, gyros, and / or other sensors disclosed herein, which can be, for example, operatively coupled to the frame 434 or otherwise attached to the user 440; and / or b) data acquired and / or processed using the remote processing module 452 and / or the remote data repository 454, including data related to virtual content, possibly transmitted to the display 432 after such processing or retrieval. The local processing and data module can be operatively coupled by communication links 438, such as via a wired or wireless communication links, to the remote processing and data module 450, which can include a remote processing module 452, a remote data repository 454, and a battery 460. The remote processing module 452 and the remote data repository 454 can be coupled by communication links 456 and 458 to the remote processing and data module 450, such that these remote modules are operatively coupled to each other and available as resources to the remote processing and data module 450. In some embodiments, the remote processing and data module 450 can include one or more of the following: image capture devices, microphones, inertial measurement units, accelerometers, compasses, GPS units, radio devices, and / or gyros. In some other embodiments, one or more of these sensors can be attached to the frame 434, or can be separate structures that communicate with the remote processing and data module 450 through wired or wireless communication paths. With continued reference to Figure 4In some embodiments, the remote processing and data module 450 can include one or more processors configured to analyze and process data and / or image information, such as including one or more central processing units (CPUs), graphics processing units (GPUs), specialized processing hardware, etc. In some embodiments, the remote data repository 454 can include a digital data storage facility, which can be accessed via the Internet or other networks in the form of a “cloud” resource configuration. In some embodiments, the remote data repository 454 can include one or more remote servers that provide information to the local processing and data module and / or the remote processing and data module 450 (e.g., information for generating augmented reality content). In some embodiments, all data is stored locally, all computing is performed locally, thereby allowing the remote modules to be used completely autonomously. Optionally, an external system including CPUs, GPUs, etc. (e.g., a system of one or more processors, one or more computers) can perform at least a portion of the processing (e.g., generating image information, processing data) and provide information to and receive information from the illustrated modules, e.g., through a wireless or wired connection. Figure 5 A perspective view of a wearable device 500 is shown, in accordance with an embodiment of the application. The wearable device 500 includes a frame 502 that is configured to support one or more projectors 504 at different locations along the frame 502 facing an interior surface. In some embodiments, the projector 504 can be attached at a location proximate to a temple 506. Alternatively, or in addition, another projector can be placed at location 508. Such a projector can include, for example, one or more liquid crystal on silicon (LCoS) modules, micro-LED displays, or fiber scanning devices, or work in conjunction therewith. In some embodiments, light from the projector 504 or from a projector disposed at location 508 can be directed into an eyepiece 510 for display to a user’s eye. A projector placed at location 512 can be slightly smaller, as this brings the projector in close proximity to the waveguide system. The closer distance can reduce the amount of light lost as the waveguide system directs light from the projector to the eyepiece 510. In some embodiments, the projector at location 512 can be used in conjunction with the projector 504 or the projector disposed at location 508. Although not depicted, in some embodiments, a projector can also be located below the eyepiece 510. The wearable device 500 is also depicted as including sensors 514 and 516. The sensors 514 and 516 can take the form of forward-facing and side-facing optical sensors that are configured to characterize the real-world environment surrounding the wearable device 500. Embodiments of the application utilize an eye tracking system to determine a user’s eye gaze position and utilize that eye gaze position for image compression processing. Reference is made to Figure 5The eye tracking camera 505 is located on the frame 502 and can be used to track the eye gaze position of a user using the wearable device 500. In other embodiments, other eye tracking systems can be used to determine the eye gaze position, and Figure 5 The eye tracking camera 505 shown in FIG. 5 is merely exemplary. As described in greater detail herein, the image compression process used in compressing and decompressing virtual content for storage in memory, internal communication, and display, among other functions, can be modified according to the eye gaze position, e.g., a higher quality compression process can be used to compress the portion of the image or video stream corresponding to the eye gaze position compared to other portions of the image or video stream located farther away from the eye gaze position. Since these farther portions of the image or video stream are located in the user's peripheral field of view, any impact of the reduction in compression quality on the user experience can be less than the benefits gained in terms of memory and processing efficiency and / or requirements. Those skilled in the art will be able to understand many variants, modifications, and alternatives. In conventional systems, image compression (e.g., JPEG compression) is implemented at a fixed quality for image or video streams that do not take into account human gaze. Since MPEG is a derivative of JPEG, embodiments of the present invention can be applied to MPEG compression processes as appropriate. By knowing the current position of human gaze and taking human gaze into account, embodiments of the present invention can reduce the quality (i.e., bandwidth) of the image in locations that the user is not gazing at (i.e., locations in the image that are spatially separated from the eye gaze position), thereby reducing the image quality in these areas and reducing the overall need to send content that the human eye cannot distinguish at a higher quality setting since the human eye is not currently focused on these non-gaze locations. Thus, embodiments of the present invention provide a video compression algorithm that takes into account human gaze and create a gaze point compression algorithm that is dependent on human gaze. In some embodiments, the JPEG algorithm receives an image and divides it into macroblocks (e.g., 16 pixels x 16 pixels). These macroblocks are then subjected to a Discrete Cosine Transform (DCT) process. The DCT process generates a set of coefficients, which are filtered so that high frequency values are eliminated (this is the key to the quality step). After this process occurs, the blocks are run-length encoded. Encoder-based gaze point (Foveation) mapping
[0003] Table 1 shows a matrix of an 8x8 pixel sub-image block according to embodiments of the present invention. An 8x8 pixel sub-image block can also be referred to as a macroblock or a tile. The 8x8 pixels are represented by the pixel values shown in the matrix. 52 55 61 66 70 61 64 73 63 59 55 90 109 85 69 72 62 59 68 113 144 104 66 73 63 58 71 122 154 106 70 69 67 61 68 104 126 88 68 70 79 65 60 70 77 68 58 75 85 71 64 59 55 61 65 83 87 79 69 68 65 76 78 94 Table 1 Table 2 is a matrix showing an example of an encoded 8x8 FDCT block according to an embodiment of the application. In a conventional system, JPEG / MPEG compression processes the entire image at a fixed quality. The filtering process results in the generation of zero data as shown in the quantized DCT block shown in Table 3. This filtering occurs at a given quality setting. As shown in Table 2, the magnitude of the values generally decreases from the upper left portion of the matrix to the lower right portion. -415 -30 -61 27 56 -20 -2 0 4 -22 -61 10 13 -7 -9 5 -47 7 77 -25 -29 10 5 -6 -49 12 34 -15 -10 6 2 2 12 -7 -13 -4 -2 2 -3 3 -8 3 2 -6 -2 1 4 2 -1 0 0 -2 -1 -3 4 -1 0 0 -1 -4 -1 0 1 2 Table 2 Table 3 is a matrix showing an example of a quantized DCT block according to an embodiment of the application. In Table 3, quantization results in a significant number of values being reduced to zero. -26 -3 -6 2 2 -1 0 0 0 -2 -4 1 1 0 0 0 -3 1 5 -1 -1 0 0 0 -4 1 2 -1 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 Table 3 Figure 6 is a schematic diagram showing the run-length encoding of a quantized DCT block according to an embodiment of the application. To encode a quantized DCT block, the run-length encoding process begins at the upper left pixel and progresses to the lower right pixel. Referring to Figure 6 , pixel 610 is encoded first, followed by pixel 612. Next, the encoding progresses to the next two rows of pixels, ending with pixel 614 and pixel 616 being encoded. The subsequent encoding process encodes pixels 618, 620, and 622. At this stage, the encoding process reverses direction, encoding pixels 624, 626, and 628. This encoding pattern is then continued until all pixels in the block have been encoded. Figure 7 is a schematic diagram of a JPEG header structure. As shown in Figure 7 , in the JPEG header structure, the default quality for the entire image is stored in the quantization table map area. Thus, a single quality setting is used to compress the entire image. As described herein, the quantization table can be applied to non-foveal areas, providing a 100% quality setting for areas corresponding to the location of the eye gaze, or the quantization table can be applied to areas further from the location of the eye gaze, providing a reduced quality setting for the foveal area. Those skilled in the art will appreciate many variations, modifications, and alternatives. Referring to Figure 7 , the segments include the start of the image, application 0 (default header), definition of quantization table (for luminance), definition of quantization table (for chrominance), start of frame, definition of Huffman table 1, definition of Huffman table 2, definition of Huffman table 3, definition of Huffman table 4, start of scan, image data (entropy encoded segment), and end of the image. The fields and values of these segments are shown in Table 4. Table 4 Embodiments of the present application maintain high quality on image blocks that are focused by the eye, while reducing the quality setting on image blocks that are not focused by the eye. These different quality settings are stored in a foveation map. The foveation map can thus be passed to the compression engine. In turn, the compression engine can selectively change the predetermined video blocks corresponding to the eye gaze location in order to compress these predetermined video blocks at high quality, while compressing other blocks at low quality. The foveation map can be created based on eye gaze information, i.e. the current focus or viewing location of the human eye can be actively determined. In embodiments of the present application, the foveation map is provided to the encoder and passed to the decoder. An additional benefit provided by embodiments of the present application is that by using the concept of video blocks, blocks with zero data (i.e. blocks that are all black) will consume less storage space or power during video display. Embodiments of the present application thus exploit a modified video block compression algorithm to achieve variable quality per block. Decoder-based foveation map The decoder can use the current DCT coefficient stream passed to it, which is included as part of the compression standard. Thus, some blocks will have more coefficients and some blocks will have fewer coefficients. However, the foveation map can be sent or passed to the decoder so that the decoder can use the location of the blocks / picture blocks of reduced quality. The decoder can thus use the foveation map to apply the required quality setting to each picture block / block. Furthermore, this information can be used to apply post-processing image filtering to remove JPEG low quality artifacts. Map implementation It is noted that a particular implementation can have an inferred 100% quality and use a global table as a fallback table, and vice versa. Embodiments of the present application can exploit a number of mechanisms to implement the selection of the quality map. As described herein, embodiments use more than one quality setting per picture, where the quality setting is defined on a per picture block basis. The foveation map provided to the encoder (e.g. a JPEG encoder) thus enables the encoder to determine which quality setting to use for a given picture block. In addition to two maps, three or more maps can be used. The foveation index (0, 1, 2..) for each block will indicate which map the encoder is to implement. Thus, we can set a range of 100%, 75%, 50%, 25% etc. quality settings. Figure 8This is a line graph of an image compressed using a single quality setting. In this example, all pixels in the image are compressed using the conventional process of applying a single quality setting to each pixel. While this process achieves uniform image compression across the entire image, the inventors have determined that processing and storage requirements can be reduced if portions of the image farther from where the user is viewing are compressed at a lower quality compared to the portion of the image corresponding to where the user is viewing, while still achieving the desired user experience. Figure 9 This is a line graph of a fixation point image having three fixation point regions according to an embodiment of the present invention. Figure 9 The image is divided into multiple regions based on eye gaze position. In this case, the user gazes at the center of the image, resulting in the eye gaze position being located at the center of the image. As described herein, a combination of... Figure 5 and Figure 22 The discussed eye-tracking system determines the eye's gaze position. Therefore, the image can be divided into a central region corresponding to the eye's gaze position and a peripheral region farther away from the eye's gaze position. In some embodiments, a gaze point map is created based on the eye's gaze position, where portions of the image closer to the eye's gaze position are mapped to a high-quality setting, while portions of the image farther away from the eye's gaze position are mapped to a lower-quality setting. Figure 9 In this model, the fixation map takes the form of two peripheral regions with lower quality settings and a central region with higher (e.g., 100%) quality settings. exist Figure 9 In the image shown, region 910 (corresponding to the left quarter of the image, i.e., the left 1 / 4) has been compressed using the first quality setting. Furthermore, region 930 (corresponding to the right quarter of the image, i.e., the right 1 / 4) has also been compressed using the first quality setting. However, region 920 (corresponding to the middle half of the image, i.e., the center 2 / 4) has been compressed using a second quality setting higher than the first quality setting. This method of dividing the image into multiple parts can be called a three-region division: the left quarter (e.g., the fixation point at 70% quality setting), the center half (e.g., the non-fixation point at 100% quality setting), and the right quarter (e.g., the fixation point at 70% quality setting). although Figure 9 The illustration shows the viewpoint map divided into three regions, but the invention is not limited to this embodiment; the image can be divided in other ways. By dividing the image into multiple regions, the quality settings for each block or tile (e.g., an 8x8 pixel block for JPEG compression) within each region can be set to predetermined quality settings for each block. Therefore, in Figure 9In the example shown, all blocks in each region are assigned the same quality setting, i.e., blocks in region 910 are assigned a first quality setting (e.g., 70%), blocks in region 920 are assigned a second quality setting (e.g., 100%), and blocks in region 930 are assigned the first quality setting (e.g., 70%), but this is not required, and individual blocks in a region can be assigned different quality settings. Thus, a gaze point map can compress a region of an image with a higher quality setting than a region farther from the eye gaze location. In some embodiments, a gaze point map can be defined as a set of regions, each region having a different quality setting assigned to blocks in the region. In some embodiments, a gaze point map can be defined as a set of regions, each region having a different quality setting assigned to blocks in the region, and a center region having a uniform quality setting assigned to blocks in the center region. In some embodiments, a gaze point map can be defined as a set of regions, each region having a uniform quality setting assigned to blocks in the region, and a center region having a different quality setting assigned to blocks in the center region. In some embodiments, a gaze point map can be defined as a set of regions, each region having a different quality setting assigned to blocks in the region, and a center region having a different quality setting assigned to blocks in the center region. Those skilled in the art will understand many variations, modifications, and alternatives. Figure 9 The three-region division shown is more complex than the two-region division. In some embodiments, a gaze point map can be defined as a set of regions, blocks in a peripheral region having a quality setting dependent on its distance from the eye gaze location, and blocks in a center region having a uniform quality setting. In other embodiments, a gaze point map can be defined as a set of regions, blocks in a peripheral region having a uniform quality setting, and blocks in a center region having a quality setting dependent on its distance from the eye gaze location. Those skilled in the art will understand many variations, modifications, and alternatives. In Figure 9 In the three-region gaze point map shown, the image / memory size is reduced by about 67% overall, while maintaining 100% quality in region 920 (i.e., the non-gaze point partition). As noted above, the non-gaze point region (i.e., the region that is not compressed or compressed using a lossless compression algorithm) can be any region identified in the gaze point map. Thus, in some embodiments, a gaze point map can be defined as a set of regions, each region having a different quality setting assigned to blocks in the region, and a center region having a uniform quality setting assigned to blocks in the center region. In some embodiments, a gaze point map can be defined as a set of regions, each region having a uniform quality setting assigned to blocks in the region, and a center region having a different quality setting assigned to blocks in the center region. In some embodiments, a gaze point map can be defined as a set of regions, each region having a different quality setting assigned to blocks in the region, and a center region having a different quality setting assigned to blocks in the center region. Those skilled in the art will understand many variations, modifications, and alternatives. Figure 9 The three-region division shown is merely an example. It is noted that if the eye gaze location is, for example, located on the right side of the image, the gaze point map can compress the right side of the image using a higher quality setting and compress the left side of the image using a lower quality setting. Thus, in the present example, if the eye gaze location is located within region 930, regions 910 and 920 will be compressed using the first quality setting, and region 930 will be compressed using a second quality setting higher than the first quality setting. In some embodiments, for example, if the eye gaze location is located within region 930, region 930 can be compressed using a higher quality setting (e.g., lossless compression), region 920 can be compressed using a medium quality setting lower than the higher quality setting, and region 910 can be compressed using a lowest quality setting lower than the medium quality setting. Thus, the gaze point of the image is a function of the eye gaze location, compressing or encoding regions including the eye gaze location with a higher quality setting than one or more regions farther from the eye gaze location. Those skilled in the art will understand many variations, modifications, and alternatives. Furthermore, although Figure 9 A set of vertical regions is shown, but this is not required by embodiments of the present invention, and the definition of regions can be performed in other ways, including horizontally oriented regions, regions defined based on distance to the eye gaze location (e.g., a set of radially defined regions), etc. Figure 10 is a second gaze point image with post-processing in gaze point regions according to another embodiment of the present application. In Figure 9 After post-processing the image shown in Figure 11 is a gaze point 3D generated image with three gaze point regions according to yet another embodiment of the present application. In Figure 11 the definition of the regions is similar to that shown in Figure 9 and Figure 10 However, since for a 3D generated image most of the image is black, the compression ratio can be higher. Using the methods described herein, a compression ratio of 87% is achieved while maintaining 100% quality in the center of the image corresponding to the eye gaze position. In this example, region 1120 is compressed using a 100% quality setting (non-gaze point is 100% quality setting), while regions 1110 and 1130 are compressed using a lower quality setting (gaze point is 20% quality setting). Since for many instances of virtual content, the image content is highest near the eye gaze position, while the peripheral regions are dark or black, embodiments of the present application are particularly suitable for virtual reality and augmented reality implementations. In some examples, all regions of the image can be compressed using a lower quality setting, and the non-gaze point regions can be compressed using a higher quality setting. For example, Figure 9 regions 910, 920, and 930 can all be compressed using the low quality setting for gaze point regions. Region 920 can also be compressed using a high quality setting. When decoding the compressed image (e.g., to reconstruct for display to a user), it can be desirable to decode the various partitions of the image in parallel. Thus, two decoders can be used to decode the compressed image. During reconstruction of the image, the decoded region 920 using the high quality setting can be superimposed on the decoded regions 910, 920, 930 (i.e., the entire image) using the low quality setting. The encoding can be JPEG (e.g., using the quality settings described above), or a technique including DSC or VDC-X (e.g., using the compression ratio) as will be discussed in more detail herein. Figure 12 is a line drawing showing an image that can be used in conjunction with multiple gaze point maps according to embodiments of the present application. In Figure 12 the image shown includes a person 1206 in partition 1210, trees 1202 in partitions 1220, 1222, 1230, 1232, and a house 1204 in partitions 1224, 1226, 1238, and 1240. Depending on the eye gaze position, different gaze point maps can be created based on the image. If the user eye gaze position is in one of the partitions 1220, 1222, 1230, or 1232, i.e., the user is looking at the trees 1202, a gaze point map can be used in which the blocks in the partitions 1220, 1222, 1230, and 1232 are compressed using a 100% quality setting (non-gaze point is 100% quality setting), while the blocks in the remaining partitions (i.e., partitions 1210, 1212, 1214, 1216, 1224, 1226, 1228, 1234, 1236, 1238, 1240, and 1242) are compressed using a lower quality setting (gaze point is 70% quality setting). Thus, compression of the image can be achieved using a gaze point map that maintains the quality of the area of the image corresponding to the eye gaze position, and the peripheral portions of the image can be compressed using a lower quality setting to conserve system resources, including memory and processing.
[0004] Alternatively, if the user eye gaze position is in one of the partitions 1224, 1226, 1238, or 1240, i.e., the user is looking at the house 1204, a gaze point map can be used in which the blocks in the partitions 1224, 1226, 1238, and 1240 are compressed using a 100% quality setting (non-gaze point is 100% quality setting), while the blocks in the remaining partitions (i.e., partitions 1210, 1212, 1214, 1216, 1220, 1222, 1228, 1230, 1232, 1234, and 1236, and 1242) are compressed using a lower quality setting (gaze point is 70% quality setting). Finally, if the user's eye gaze position is in partition 1210, i.e., the user is looking at person 1206, then a gaze point map can be used in which blocks in partition 1210 are compressed using a 100% quality setting (non-gaze point is 100% quality setting) and blocks in the remaining partitions (i.e., partitions 1212, 1214, 1216, 1220, 1222, 1224, 1226, 1228, 1230, 1232, 1234, and 1236, 1238, 1240, and 1242) are compressed using a lower quality setting (gaze point is 70% quality setting). In some embodiments, the quality setting used for the remaining partitions varies, e.g., as a function of distance from the eye gaze position. In these embodiments, blocks in partitions 1212, 1214, and 1216 can be compressed using a 90% quality setting, blocks in partitions 1220, 1222, 1224, 1226, and 1228 can be compressed using an 80% quality setting, and blocks in partitions 1230, 1232, 1234, 1236, 1238, 1240, and 1242 can be compressed using a 70% quality setting. In some examples, partitions 1210-1242 can be compressed using a technique including DSC or VDC-X (e.g., using a compression ratio) rather than being encoded using JPEG (e.g., with the quality settings described above). For example, based on the eye gaze position, a non-tile-based compression technique like DSC can be used to compress partitions close to the eye gaze position with a lower compression ratio, while compressing partitions farther from the eye gaze position with a higher compression ratio. Figure 13 A simplified flowchart of a method of compressing an image according to an embodiment of the present application is shown. The method 1300 includes receiving an image (1310), determining an eye gaze position of a user (1312), and generating a gaze point map based on the eye gaze position (1314). The image can be an image included in a video stream. Determining an eye gaze position of a user can utilize an eye tracking system that provides the eye gaze position as a function of time. The gaze point map defines a compression quality of blocks and varies according to position in the image, with blocks in areas close to the eye gaze position being compressed using a higher quality setting and blocks in areas farther from the eye gaze position being compressed using a lower quality setting. In some embodiments, the gaze point map includes a first region of the image and a second region of the image. Figure 9 In the example shown, the gaze point map includes three regions, but the present application is not limited to this particular implementation and two regions or more than three regions can be defined. Further, blocks in a given region can be compressed using a uniform quality setting or can be compressed using different quality settings according to a particular implementation. In some embodiments, the gaze point map includes a first region of the image and a second region of the image. The method also includes compressing a first region of the image using a first quality setting and compressing a second region of the image using a second quality setting (1316). In some embodiments, the first quality setting is an uncompressed quality setting or a lossless compression quality setting. Therefore, blocks in the first region are compressed at a higher quality than other parts of the image. The second quality setting is a lower quality setting, such as 70% of the quality setting, which reduces the data corresponding to the compressed image in these regions. As mentioned above, since these regions are located in the user's peripheral field of view due to the user's eye gaze, any quality loss is offset by savings in memory and processor usage. The data compression processes for the first and second regions can be performed sequentially or in parallel, depending on the specific application. Compressed images or videos (which may be called foveated images or videos) can be mapped to foveated points. Figure 1 It can be transmitted to the display system (1318), or it can be mapped to the gaze point. Figure 1 It is stored in memory (1319). Compressed images or videos along with gaze mapping Figure 1 In embodiments where the image is stored in memory, method 1300 includes retrieving a gaze point image and a gaze point map from memory (1320), setting a first region of the decompressed image using a first quality setting, and setting a second region of the decompressed image using a second quality setting (1340). The compressed image or video, along with the gaze point map... Figure 1 In an embodiment where the image is transmitted to a display system, method 1300 includes receiving a foveation image and a foveation map (1320), setting a first region of the decompressed image using a first quality setting, and setting a second region of the decompressed image using a second quality setting (1340). The decompression processes for the first and second regions can be performed sequentially or in parallel, depending on the specific application. The two regions can be merged to form a final image suitable for display (1342). The final image is then displayed on a display device (1344). It should be understood that Figure 13 The specific steps illustrated provide a particular method for compressing images according to embodiments of the present invention. According to alternative embodiments, other sequences of steps may also be performed. For example, alternative embodiments of the present invention may perform the above steps in a different order. Furthermore, Figure 13 The steps shown may include multiple sub-steps, which can be performed in various orders as needed. Furthermore, other steps may be added or removed depending on the specific application. Many variations, modifications, and alternatives will be understood by those skilled in the art. Figure 14 This is a simplified schematic diagram illustrating a gaze-based image gaze point system according to an embodiment of the present invention. (Reference) Figure 14The gaze-based image foveation system 1400 includes a wearable device 1410 (e.g., a wearable device containing an ASIC performing the illustrated operations) that receives images or videos suitable for display to a user. One or more communication interfaces 1420 can be used to receive the images or videos. In the illustrated embodiment, WiFi, USB, DisplayPort (DP) or other communication protocols are used to receive the image or video content. In this embodiment, the uncompressed content is MPEG video. The wearable device 1410 also receives eye gaze information from the eye tracking system 1405. The eye tracking system 1405 may include one or more sensors suitable for measuring eye position and orientation, and may provide data that the eye gaze processor 1430 can use to calculate the user's eye gaze. Figure 14 In the illustrated embodiment, the eye gaze processor 1430 is implemented using a CPU or neural processing unit (NPU) controller, but other processors may also be used. Many variations, modifications, and alternatives will be understood by those skilled in the art. like Figure 14 As shown, in some embodiments, an image or video is passed to an image compression processor 1422, which performs a process of forming a compressed image / video (e.g., a foveated image / video) based on the user's eye gaze, as described in more detail herein. Different foveated processing methods may be used depending on the specific application, including tile-based foveated processing (e.g., JPEG or DSC foveated processing, as described in more detail herein), sparsity-based compression processing, etc. In some embodiments, for example, if the image is remotely compressed before being received by one or more communication interfaces 1420, the image compression processor 1422 is bypassed, and the image or video is passed to memory 1424 for storage. When an image or video compressed using image compression processor 1422 or remotely compressed is retrieved from memory 1424, the image decompression process can be performed using decompression processor 1426 and eye gaze information provided by eye gaze processor 1430. In embodiments where the image is remotely compressed and bypasses image compression processor 1422, decompression processor 1426 can decode the compressed image. The original or reconstructed image is then passed to warp / depth reprojection processor 1428. After warping or depth re-projection, the data provided by the eye gaze processor 1430 can again be utilized to compress the warped image using a variable quality encoder 1432 containing processor component 1431 that represents the image gaze point based on the eye gaze location. Different gaze point processing can be used depending on the particular application, including tile-based gaze point processing, sparsity-based compression processing, etc. In some embodiments, the variable quality encoder 1432 containing processor component 1431 is bypassed. As described above, the variable quality encoder 1432 can perform a JPEG encoding process to form a gaze point image based on the eye gaze, where the quality of the image varies across the image, providing high quality in the image regions corresponding to the user's eye gaze, and lower quality in image regions further away from the eye gaze location. Thus, a reduced size gaze point encoded as well as sparsity encoded image can be formed, while maintaining the required image quality. The encoded image is then provided to a Mobile Interface Processor Interface (MIPI) device 1434 for subsequent transmission to a display system. The MIPI device 1434 of the wearable device 1410 can be connected to a MIPI device 1442 of a display system 1440 that includes a variable quality decoder 1444 containing processor component 1443 that performs de-gaze point processing based on the eye gaze location, as well as a display device 1446, such as an LCOS display or a micro light emitting diode (pLED) display. As shown in the implementation of the variable quality decoder 1444, JPEG / DSC tile-based encoded data or N-way compressed encoded data (e.g., N-way DSC) can be received in a first communication channel, and a quality map (Q-map) (e.g., a gaze point map) can be received in a second communication channel for use in the decoding process. Alternatively, the Q-map can be received using an embedded line format or other suitable format. Figure 14 As shown in the implementation of the variable quality decoder 1444, JPEG / DSC tile-based encoded data or N-way compressed encoded data (e.g., N-way DSC) can be received in a first communication channel, and a quality map (Q-map) (e.g., a gaze point map) can be received in a second communication channel for use in the decoding process. Alternatively, the Q-map can be received using an embedded line format or other suitable format. As shown, a JPEG decoding process can be performed by the variable quality decoder 1444 containing processor component 1443 to form a final image based on the gaze point image generated by the variable quality encoder 1432 containing processor component 1431. Thus, embodiments of the present application reduce system memory and transmission requirements, e.g., reducing the amount of data transmitted between MIPI devices, while maintaining the required image quality. The decoded image will then be displayed using the display device 1446. Figure 14 As shown, a JPEG decoding process can be performed by the variable quality decoder 1444 containing processor component 1443 to form a final image based on the gaze point image generated by the variable quality encoder 1432 containing processor component 1431. Thus, embodiments of the present application reduce system memory and transmission requirements, e.g., reducing the amount of data transmitted between MIPI devices, while maintaining the required image quality. The decoded image will then be displayed using the display device 1446. In certain embodiments, the variable quality encoder 1432 is bypassed, and the warped image is transmitted to the display system 1440 using the MIPI device 1434 without variable quality image compression. In these embodiments, the variable quality decoder 1444 is also bypassed. Although the above embodiments employ a tile-based (also referred to as block-based) JPEG compression algorithm, embodiments of the present application are not limited to this particular compression standard, and other compression standards can be used in conjunction with various embodiments of the present application. For example, Figure 15 to Figure 21 Techniques are described for compressing video data using run-length encoding in conjunction with DSC and VDC-X. Figure 15 A plot of compression levels obtained as a function of time for a sparse compression system implementation and a DSC-SPARSE system implementation according to embodiments of the present application is shown in relation to consecutive frames and frequency. In this example, the compression levels are obtained for a video sequence of 10,000 frames. Figure 15 In this example, an alternating algorithm is implemented in which either a mask-based compression method or DSC is used for compression of each frame, depending on whether the frame is implemented based on a mask-based compression method or a full frame fixed compression (e.g., DSC). As shown in Figure 15 , each frame is analyzed and the number of rows having pixels with a luminance level less than a threshold value is determined. If the mask-based compression method results in a compression level greater than a compression threshold (e.g., 37%), then the frame is compressed using the mask-based compression method. In Figure 15 , this results in the first approximately 3800 frames being compressed using the mask-based compression method. If the mask-based compression method results in a compression level of the compressed frame that is less than 37%, e.g., a frame with very little black content, then the DSC method is used. This results in a compression value of 37% for these frames. Referring to Figure 15 , frames represented by a blue compression value less than 37% are compressed using DSC, effectively setting a minimum compression to 37%. Thus, the compression value for frames in sets A and B is 37%, rather than a lower value that would be achieved using the mask-based compression method. Figure 16 A plot of the number of frames versus compression ratio for a sparse compression system implementation and a DSC-SPARSE system implementation according to embodiments of the present application is shown in Figure 16 . As shown, the number of frames with a compression ratio less than approximately 37% is reduced to zero because the mask-based compression method is used for frames that can be compressed at a level greater than 37%, or the frame-based compression method (e.g., DSC) is used for the remaining frames that cannot be compressed at a level greater than 37% using the mask-based compression method. Thus, although the mask-based compression method alone would result in multiple frames with a compression ratio less than 37%, the alternating method provided by embodiments of the present application limits the minimum compression level to approximately 37%, as shown in Figure 16The mask-based compression method can provide a high level of compression for frames containing a large amount of black pixel content; while the frame-based compression method establishes a lower limit for the level of compression, e.g., 37% in the present embodiment, for frames containing limited black pixel content. Those skilled in the art will appreciate that the minimum level of compression need not be 37%, which is merely an example, and other minimum levels of compression can be employed depending on the particular application. Those skilled in the art will appreciate that there are many variations, modifications, and alternatives. Information regarding the compression method used for each frame can be provided to the endpoint (e.g., decoder or display) so that the endpoint uses the appropriate decompression method when reconstructing each frame. Figure 17 A simplified flowchart of a method for compressing image frames using an alternating compression algorithm according to embodiments of the present application. The method 1700 includes receiving a frame of video data (1710). The method also includes determining a number of rows of groups of pixels in the frame having a luminance level less than a threshold value (1712). If the number of rows is greater than or equal to a compression threshold value (1714), the frame is compressed using a mask-based compression method (1720). If the number of rows is less than the compression threshold value, the frame is compressed using a frame-based compression method (1722). If there are other frames (1730), the method operates on the next frame of video data by receiving a frame of video data (1710). Otherwise, the method ends (1740). Thus, embodiments of the present application alternate between different compression methods for each frame depending on the level of compression achievable by each compression method. It should be understood that Figure 17 The specific steps shown provide one particular method for compressing image frames using an alternating compression algorithm according to embodiments of the present application. Other sequences of steps can be performed according to alternative embodiments. For example, alternative embodiments of the present application can perform the steps described above in a different order. Moreover, Figure 17 Each of the steps shown can include a number of sub-steps that can be performed in various orders depending on the needs of each step. Moreover, other steps can be added or removed depending on the particular application. Those skilled in the art will appreciate many variations, modifications, and alternatives. According to some embodiments of the present application, there will be an embedded image line control or alternate control mechanism that will provide information to the endpoint display on a frame-by-frame basis regarding which system to use to decode incoming MIPI frames. In addition, a dummy MIPI lane can be utilized to indicate the compression ratio used by the endpoint display. Some embodiments of the present invention change the compression quality based on eye tracking, providing higher compression ratios for the fovea region at the expense of quality loss. It does this for the MIPI interface, reducing the amount of data sent over MIPI to the LCOS / uLED display. As a result, these embodiments also save power consumption. Embodiments of the present invention reduce the amount of stream-based data sent over MIPI compression. In addition, embodiments change the compression quality based on eye tracking, providing higher compression ratios for the fovea region at the expense of quality loss. Furthermore, embodiments allow for higher compression ratios for stream-based compression techniques and allow for quality to be maintained for regions observed by the user. As a result, embodiments allow for higher compression ratios while maintaining quality. For stream-based compression standards, such as DSC and VESA Display Compression (VDC-X), a low latency implementation is employed. With this low latency reaction, the spatial WARP adjustments made previously are still applicable.
[0005] Figure 18 is a simplified diagram showing an image frame divided into a high quality region and a low quality region according to embodiments of the present invention. Figure 18 The illustrated image 1800 includes a high quality region 1810 and a low quality region 1820. As discussed in greater detail below, the high quality region 1810 will be compressed and decompressed using a first quality setting or compression level, while the low quality region 1820 (or the entire image) will be compressed and decompressed using a second quality setting or compression level, saving memory and bringing other benefits. For example, a single decoder can be used by not compressing the high quality region 1810 and compressing the low quality region using a single decoder. If the high quality region 1810 is small relative to the entire image, significant savings can be achieved. Additional description related to varying the size of the high quality region is provided in U.S. Provisional Patent Application No. 63 / 543,876, filed October 12, 2023, the disclosure of which is hereby incorporated by reference in its entirety herein for all purposes. DSC Conventional DSC does not provide variable quality compression. Instead, DSC employs 24-bit color encoding and compresses it to 15 / 12 / 10 / 8 bits. The higher the compression ratio (24→8 bpp), the worse the impact on quality. As for the quality required for eye focus partitioning, embodiments are able to maintain a PSNR quality setting of, for example, 60 dB or more as described above. According to Figure 6The use case analysis shown, the inventors have determined that this occurs only at 37% compression configuration (24→ 15bpp). However, in practice, only the area where the eye is currently focused will use this compression setting. The outer gaze point area (e.g., the part of the image that is further away from the eye gaze position) can sustain a lower quality, e.g., 75% compression level (24→ 8bpp). Thus, for a neighborhood-based compression standard like DSC, embodiments will divide the home screen into high quality areas and low quality areas (as shown in Figure 18 ) or smaller partitions (as shown in Figure 20 ), each with a different compression ratio. The compression ratio chosen will depend on the current eye gaze position. Thus, referring to Figure 18 where the eye gaze position is within a high quality area 1810, the high quality area 1810 can be compressed at a lower compression level (e.g., 24→ 15bpp) while the low quality area 1820 can be compressed at a higher compression level (e.g., 24→ 8bpp). In some examples, the low quality area 1820 can be compressed at an even higher compression level (e.g., 24→ 6bpp). In embodiments that compress the entire image using a higher compression level (as described in more detail herein), the high quality area 1810 can be overlaid on the entire image when the image is reconstructed. Figure 19 is a simplified flowchart showing a method 1900 of compressing an image using different compression ratios for high quality areas and low quality areas, according to embodiments of the present application. The method 1900 includes determining an eye gaze position of a user (1910), generating a gaze map containing a first area of the image and a second area of the image (1912), and compressing the first area using a first compression ratio and the second area using a second compression ratio (1914). The image can be an image contained in a video stream. Determining an eye gaze position of a user can utilize an eye tracking system that provides the change in eye gaze position over time. The gaze map defines the compression ratio at which portions of the image are compressed and varies according to the position in the image relative to the eye gaze position, with areas closer to the eye gaze position being compressed using a lower compression ratio and areas further away from the eye gaze position being compressed using a higher compression ratio. In Figure 18In the illustrated example, the gaze point map contains two regions, but the application is not limited to this particular implementation, and three or more regions can be defined. In some embodiments, the gaze point map includes a first region of the image and a second region of the image. The method 1900 can be referred to as N-way compression (e.g., DSC, VDC-X, or JPEG), where N refers to the number of regions determined for the image. For example, depending on the eye gaze location, a high quality region of the image, a medium quality region around the high quality region, and a low quality region can be determined. The techniques of the method 1900 can then be used as a three-way compression, with different compression ratios used for each region. Referring back to Figure 18 In some examples, the low quality region 1820 can encompass the entire image, including the portion of the image in the high quality region 1810 that is characterized by the eye gaze location. When decoding the compressed image (e.g., to reconstruct for display to the user), it can be desirable to decode the various partitions of the image in parallel. For example, as discussed above, the high quality region 1810 can be decoded using a first DSC decoder, and the low quality region 1820 can be decoded using a second DSC decoder. Figure 18 For an image divided into a high quality region 1810 and a low quality region 1820 as illustrated, the low quality region 1820 can be considered the entire image. For example, for a 2 kilopixel x 2 kilopixel (4 million pixels total) image, the low quality region 1820 can be the entire 4 million pixel image, and can be compressed using a high compression level (e.g., 24→8 bpp). The high quality region 1810 can be determined based on the current eye gaze location, and for example, can be a 1 kilopixel x 1 kilopixel region (1 million pixels total). The high quality region 1810 can be compressed using a low compression level (e.g., 24→15 bpp). Thus, the compressed image can be decoded using two DSC decoders. During the image reconstruction process, the decoded high quality region can be superimposed on the decoded low quality region. Figure 20 is a simplified diagram illustrating an image frame divided into a high quality partition and a low quality partition, according to an embodiment of the application. As discussed in more detail below, Figure 20 The partitioned image frame 2000 illustrated in FIG. 2 can be used to define a gaze point map that defines how different partitions of the image are compressed with what compression ratio, such that the compression ratio or other compression quality metric varies as a function of location in the image relative to the eye gaze location. For example, a partition close to the eye gaze location can be compressed using a lower compression ratio, while a partition farther from the eye gaze location can be compressed using a higher compression ratio. Referring to Figure 20The four partitions 2010, 2012, 2014, and 2016 containing the high quality region 2002 (i.e., the region corresponding to the current eye gaze position) will be compressed at a lower compression level (e.g., 24→15 bpp), while the remaining partitions (which can be referred to as peripheral partitions or low quality partitions) will be compressed at a higher compression level (e.g., 24→8 bpp). Thus, when the compressed image is reconstructed for display to the user, the quality of the high quality region corresponding to the eye gaze position will be higher than the remaining portions of the image that are further away from the eye gaze position. Thus, embodiments of the present application provide a gaze point image based on the eye gaze position and reduce storage and transmission requirements. In Figure 20 In some embodiments of the illustrated example, all of the partitions 2010-2046 of the image can be compressed at a high compression ratio (e.g., 24→8 bpp). The four partitions 2010, 2012, 2014, and 2016 containing the high quality region can also be compressed at a lower compression ratio (e.g., 24→15 bpp). Using a decoder, all of the partitions 2010-2046 compressed at the high compression ratio can be decoded at the higher compression ratio, while the four partitions 2010, 2012, 2014, and 2016 compressed at the lower compression ratio can also be decoded at the lower compression ratio. During the image reconstruction process, the decoded high quality partitions 2010, 2012, 2014, and 2016 can be superimposed on the decoded low quality partitions 2010-2046. In some embodiments, the gaze point map can define the partitions that coincide with the high quality region. For example, the partitions 2010-2016 can contain only the high quality region characterized by the eye gaze position and not portions of the image in the low quality region. As with N-way compression, in the partition-based DSC technique, multiple DSC decoders can be required to decode the compressed image. For example, four DSC decoders can be used to decode the compressed image, where one decoder is used to decode the high quality partitions 2010-2016, another decoder is used to decode the partitions 2020-2026, a third decoder is used to decode the partitions 2030-2036, and a fourth decoder is used to decode the partitions 2040-2046, each decoder using a compression ratio for each set of partitions based on proximity to the eye gaze position. In some embodiments, depending on the memory capacity (e.g., SRAM) of the system used for decoding, a single decoder can be implemented with acceptable latency when decoding the compressed image. The image can be an image contained in a video stream. Determining the eye gaze position of the user can utilize an eye tracking system that provides the change in eye gaze position over time. The gaze point map defines the compression ratio at which different zones of the image (e.g., zones 2010-2016, zones 2020-2026, zones 2030-2036, and zones 2040-2046) are compressed, and varies according to the location in the image relative to the eye gaze position, with zones closer to the eye gaze position being compressed using a lower compression ratio, and zones farther from the eye gaze position being compressed using a higher compression ratio. In Figure 20 In the example shown, the gaze point map contains 16 zones, but the present application is not limited to this particular implementation, and more or fewer than 16 zones can be defined. The methods described herein can be referred to as zone-based compression (e.g., DSC, VDC-X, or JPEG) methods. Although only two compression levels are shown in some of the examples above, embodiments of the present application are not limited to these particular compression levels, but can utilize an additional number of compression levels. For example, zones 2010-2014 can be compressed using a compression level of 37% (i.e., 24→ 15 bpp), while zones 2020, 2022, 2024, and 2026, which are farther from the high quality region than zones 2010-2014, can be compressed using a compression level of 50% (i.e., 24→ 15 bpp), zones 2030, 2032, 2034, and 2036, which are farther from the high quality region than zones 2020-2026, can be compressed using a compression level of 58% (i.e., 24→ 12 bpp), and zones 2040, 2042, 2044, and 2046, which are farther from the high quality region than zones 2010-2016, can be compressed using a compression level of 67% (i.e., 24→ 8 bpp). Thus, the use of two compression levels is merely exemplary. Furthermore, for some zones, the compression level can be 0%, i.e., uncompressed, including the zones corresponding to the eye gaze position and the high quality region. Thus, the compressed image can have both uncompressed zones as well as compressed zones. Those skilled in the art will recognize many variations, modifications, and alternatives. Furthermore, although Figure 20 In the example shown in FIG. 2, 16 uniform region zones are shown, but this is not required, and other numbers of zones can be used, including zones of different sizes, with smaller zones adjacent to the high quality region, and larger zones (e.g., zones compressed at a higher level) farther from the high quality region. Thus, the number of compression levels, the compression levels, the number of zones, and the size of the zones can be varied as appropriate for a particular application. Those skilled in the art will recognize many variations, modifications, and alternatives. As the frame size is reduced due to image compression, the communication interface (e.g., MIPI interface) can be modified to enter a low-power data transfer mode, or even a super low-power sleep mode, thus saving computational resources and reducing power consumption. Eventually, the reconstruction of the compressed image can be performed before displaying it to the user. Figure 21 is a simplified flowchart illustrating a method 2100 of compressing an image using different compression ratios for high-quality partitions and low-quality partitions according to embodiments of the application. The method 2100 comprises determining the eye gaze position of a user (2110), generating a gaze map containing a first partition and a second partition of the image (2112), and compressing the first region using a first compression ratio and the second region using a second compression ratio (2114). It should be understood that, Figure 19 and Figure 21 The specific steps shown in Figure 19 and Figure 21 Each of the steps shown in VDC-X VDC-X compression standards (e.g., VDC-M) use a tile-based approach rather than a nearest neighbor approach. Such compression standards encode different tiles at different quality settings, but the goal of such conventional compression is to maintain a constant frame size (i.e., bit rate) overall. Thus, once a compression ratio is selected, each tile is changed to maintain a constant bit rate. In conjunction with embodiments of the application, using this compression standard, the compression of the video image is based not only on the bit rate, but also on the eye gaze position of the user. For example, the four partitions 2010, 2012, 2014, and 2016 containing the high-quality region (i.e., the region corresponding to the current eye gaze position) will be compressed at a higher quality setting than the remaining partitions (which can be referred to as peripheral partitions), which will be compressed at a lower quality setting than the partitions 2010-2016. Some embodiments of the application do not maintain a constant bit rate, such that each frame size varies over time, and the transmission interface (e.g., MIPI) is placed in a low-power mode when not in use. Similar to the DSC-based approach discussed above, for the VDC-X tile-based approach, embodiments encode the quality of each tile based on the current position of the user's eye gaze. As Figure 20As shown, eye gaze information provided by an eye gaze tracking system of the AR system is used to compress tiles according to the VDC-X standard based on the distance of the tile from the eye gaze position. Accordingly, embodiments of the application are able to vary the frame size or bit rate per frame and use the current eye gaze information to select which tile (VDC-X) or partition (DSC) has a higher quality than the gaze point region with a lower quality setting. In some embodiments, the N-way compression or partition-based compression described above can implement JPEG as the compression standard, rather than DSC or VDC-X. In these embodiments, the compression ratios for high / low quality regions and / or high / low quality partitions can instead refer to the quality settings of the JPEG standard. Figure 22 is a simplified block diagram illustrating components of an AR system according to embodiments of the application. As Figure 22 The AR system 2200 shown can be incorporated into the AR devices described herein. Figure 22 A diagrammatic representation of one embodiment of an AR system 2200 is provided that can perform some or all of the steps of the methods provided by various embodiments. It should be noted that Figure 22 only intended to provide a generalized illustration of various components, any or all of which can be utilized as appropriate. Therefore, Figure 22 various system elements will be broadly described. The AR system 2200 is shown to include hardware elements that can be electrically coupled via a bus 2205 or otherwise in communication (as appropriate) with one another. These hardware elements can include one or more processors 2210 including, without limitation, one or more general-purpose processors and / or one or more special-purpose processors such as digital signal processing chips, graphics acceleration processors, and / or the like; one or more input devices 2215 including, without limitation, a mouse, a keyboard, a camera, and / or the like; and one or more output devices 2220 including, without limitation, a display device, a printer, and / or the like. Additionally, the AR system 2200 includes an eye tracking system 2255 that can provide the AR system with the eye gaze position of a user. With the processor 2210, the gaze point image compression techniques discussed herein can be implemented. The AR system 2200 can also include and / or be in communication with one or more non-transitory storage devices 2225, which can include, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device such as a random access memory (RAM), and / or a read-only memory (ROM), which can be programmable, flash- updateable, and / or the like. Such storage devices can be configured to implement any appropriate data stores, including without limitation, various file systems, database structures, and / or the like. The AR system 2200 can also include a communication subsystem 2219, which can include without limitation a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, cellular communication The AR system 2200 can also include software elements, shown as being currently located within the working memory 2260, including an operating system 2262, device drivers, executable libraries, and / or other code, such as one or more application programs 2264, which can comprise computer programs provided by various embodiments, and / or can be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) described above can be implemented as code and / or instructions executable by a computer and / or a processor within a computer; in an aspect, such code and / or instructions can be used to configure and / or adapt a general purpose computer or other A set of these instructions and / or code might be stored on a non-transitory computer-readable storage medium, such as the storage device(s) 2225 described above. In some cases, the storage medium might be incorporated within a computer system, such as the AR system 2200. In other embodiments, the storage medium might be separate from a computer system (e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure and / or adapt a general purpose computer with the instructions / code stored thereon. These instructions might take the form of executable code, which is executable by the AR system 2200 and / or might take the form of source code, which can be compiled, or interpreted, to be executable by the AR system 2200. The instructions might be stored in object-oriented or other Those skilled in the art will appreciate that substantial variations can be made in accordance with specific requirements. For example, customized hardware might also be used, and / or particular elements might be implemented in hardware, software (including portable software, such as applets, etc.), or both. Further, connection to other computing devices such as network input / output devices can be employed. As described above, in one aspect, some embodiments can employ a computer system (such as AR system 2200) to perform methods in accordance with various embodiments of the technology. According to a set of embodiments, some or all of the procedures of such methods are performed by AR system 2200 in response to processor 2210 executing one or more sequences of instructions. The instructions can be incorporated into the operating system 2262 and / or other code of the AR system 2200 (such as an application program 2264) incorporated into the work memory 2260. Such instructions can be read into the work memory 2260 from another computer readable medium such as one or more storage device(s) 2225. Merely by way of example, execution of the sequences of instructions contained in the work memory 2260 might cause the processor(s) 2210 to perform one or more procedures of the methods described herein. The methods described herein can be The terms "machine-readable medium" and "computer-readable medium," as used as herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. In an implementation implemented using AR system 2200, various computer-readable media might be involved in providing instructions / code to processors 2210 for execution. Such a medium might take many forms, including but not limited to, non-volatile media, and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device(s) 2225. Volatile media includes dynamic memories, such as the work memory 2260. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, punchcards, papertape, any other physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave as described hereinafter, or any other medium from which a computer can read instructions and / or code. Various forms of computer-readable media can be involved in carrying one or more sequences of instructions to the processor(s) 2210 for execution. Merely by way of example, the instructions can initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer might load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received by a computing device, such as AR system 2200. These signals might then be directed to the work memory 2260 of the computing device. The communication subsystem 2219 and / or components thereof typically will receive signals, and the bus 2205 then might carry the signals and / or data, instructions, etc. received by the communication subsystem 2219 to the working memory 2260, from which the processor 2210 retrieves and executes the instructions. The working memory 2260, in turn, can be used by the processor 2210 to store and / or write various data noticed during the execution of instructions. Various examples of the present disclosure are provided below. Any reference in this document to a series of examples should be understood as a reference to each of the examples individually (e.g., "Examples 1-4" should be understood as "Example 1, 2, 3, or 4"). Example 1 is a method of compressing an image, the method comprising: determining an eye gaze position of a user; generating a gaze point map based on the eye gaze position, wherein the gaze point map comprises a first region of the image and a second region of the image; and compressing the first region of the image using a first quality setting and the second region of the image using a second quality setting. Example 2 is the method of Example 1, wherein determining the eye gaze position comprises: using an eye tracking camera of an augmented reality device. Example 3 is the method of Examples 1-2, wherein the gaze point map comprises a central region and a peripheral region.
[0006] Example 4 is the method of Examples 1-3, wherein the image comprises virtual content generated by an augmented reality device. Example 5 is the method of Examples 1-4, wherein the image is included in a virtual content video stream. Example 6 is the method of Examples 1-5, wherein compressing the first region of the image using the first quality setting comprises: compressing all blocks in the first region using the first quality setting. Example 7 is the method of Examples 1-6, wherein the first quality setting is greater than the second quality setting. Example 8 is the method of Examples 1-7, wherein the first quality setting is 100%. Example 9 is the method of Examples 1-8, further comprising: post-processing image content in at least one of the first region or the second region. Example 10 is the method of Examples 1-9, wherein the compressing results in a compressed image, the method further comprising: decoding the compressed image using the gaze point map. Example 11 is the method of examples 1-10, wherein: the first region of the image comprises a plurality of first blocks; the second region of the image comprises a plurality of second blocks; compressing the first region of the image comprises compressing each of the plurality of first blocks using the first quality setting; and compressing the second region of the image comprises compressing each of the plurality of second blocks using the second quality setting. Example 12 is the method of claim examples 1-11, further comprising: decompressing the first region of the image using the first quality setting; decompressing the second region of the image using the second quality setting; and displaying the image to the user. Example 13 is the method of examples 1-12, wherein the second region of the image comprises the first region of the image. Example 14 is the method of examples 1-13, wherein the compressing produces a compressed image, the method further comprising: decoding the compressed image using the gaze point map to produce a decoded first region and a decoded second region; and reconstructing the image by overlaying the decoded first region on the decoded second region. Example 15 is an augmented reality (AR) system, comprising: a wearable device comprising: a frame; a projector coupled with the frame; a display optically coupled with the projector; and an eye tracking system; a memory; and a processor configured to: receive an eye gaze position from the eye tracking system; generate an image; generate a gaze point map based on the eye gaze position, wherein the gaze point map comprises a first region of the image and a second region of the image; and compress the first region of the image using a first quality setting and compress the second region of the image using a second quality setting. Example 16 is the AR system of example 15, wherein the projector comprises one projector of a set of projectors, the display comprises one display of a set of displays, and the eye tracking system comprises a set of eye tracking devices. Example 17 is the AR system of examples 15-16, wherein determining the eye gaze position comprises: using an eye tracking camera of the augmented reality device. Example 18 is the AR system of examples 15-17, wherein the gaze point map comprises a central region and a peripheral region. Example 19 is the AR system of examples 15-18, wherein the image comprises virtual content generated by the augmented reality device. Example 20 is the AR system of examples 15-19, wherein the image is included in a virtual content video stream. Example 21 is the AR system of examples 15-20, wherein compressing the first region of the image using the first quality setting comprises compressing all blocks in the first region using the first quality setting. Example 22 is the AR system of examples 15-21, wherein the first quality setting is greater than the second quality setting. Example 23 is the AR system of examples 15-22, wherein the first quality setting is 100%. Example 24 is the AR system of examples 15-23, wherein the processor is further configured to post-process image content in at least one of the first region or the second region. Example 25 is the AR system of examples 15-24, wherein compressing results in a compressed image, and wherein the processor is further configured to decode the compressed image using the gaze point map. Example 26 is the AR system of examples 15-25, wherein: the first region of the image comprises a plurality of first blocks; the second region of the image comprises a plurality of second blocks; compressing the first region of the image comprises compressing each of the plurality of first blocks using the first quality setting; and compressing the second region of the image comprises compressing each of the plurality of second blocks using the second quality setting. Example 27 is the AR system of examples 15-26, wherein the processor is further configured to decompress the first region of the image using the first quality setting; decompress the second region of the image using the second quality setting; and display the image to the user. Example 28 is the AR system of examples 15-27, wherein the second region of the image comprises the first region of the image. Example 29 is the AR system of examples 15-28, wherein compressing results in a compressed image, and wherein the processor is further configured to decode the compressed image using the gaze point map to produce a decoded first region and a decoded second region; and reconstruct the image by overlaying the decoded first region on the decoded second region. Example 30 is a non-transitory computer-readable medium comprising program code executable by a processor of a user-wearable device, the program code executable by the processor to: determine an eye gaze position of a user; generate a gaze point map based on the eye gaze position, wherein the gaze point map comprises a first region of the image and a second region of the image; and compress the first region of the image using a first quality setting and compress the second region of the image using a second quality setting. In the foregoing specification, the disclosure has been described with reference to specific embodiments thereof. It is evident, however, that various modifications and changes can be made thereto without departing from the broader spirit and scope of the disclosure. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. Indeed, it is to be appreciated that the systems and methods of the present disclosure each have a plurality of novel aspects, no single one of which is solely responsible or necessary to the desired attributes offered by the disclosure. Various features described in the context of separate embodiments can also be implemented in combination with each other. Conversely, various features described in the context of a single embodiment can also be implemented on other embodiments, alone or in any suitable
[0007] It should be understood that conditional language used herein, such as “may,” “can,” “perhaps,” “may,” “for example,” etc., unless expressly stated otherwise or understood in the context, is generally intended to express that some embodiments include certain features, elements, and / or steps, while other embodiments do not include these features, elements, and / or steps. Therefore, such conditional language is not generally intended to imply that features, elements, and / or steps are necessary in any way for one or more embodiments, or that one or more embodiments necessarily include logic (whether entered or prompted by the author) for determining whether such features, elements, and / or steps are included in a particular embodiment or whether they are performed in any particular embodiment. The terms “comprising,” “including,” “having,” etc., are synonymous and used in an open-ended inclusive manner, and do not exclude other elements, features, actions, operations, etc. Furthermore, the use of the term “or” is inclusive (not exclusive), and therefore, when used to connect a series of elements, the term “or” refers to one, some, or all of the elements in the list. Additionally, unless otherwise stated, the articles “a,” “an,” and “the” used in this application and the appended claims should be understood as “one or more” or “at least one.” Similarly, while operations may be depicted in a specific order in the figures, it should be understood that these operations need not be performed in the specific order or sequence shown, nor is it necessary to perform all of the shown operations to achieve the desired result. Furthermore, the figures may schematically depict another example flow in the form of a flowchart. However, other operations not depicted may be incorporated into the schematically shown example methods and flows. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the shown operations. Furthermore, in other embodiments, these operations may be rearranged or reordered. In some cases, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated into a single software product or packaged into multiple software products. Furthermore, other embodiments are also included within the scope of the following claims. In some cases, the operations described in the claims may be performed in a different order and still achieve the desired result. Therefore, the claims are not intended to limit the embodiments shown herein, but should be given the widest scope consistent with the content of this disclosure, the principles and novel features disclosed herein. It should also be understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or alterations can be made by those skilled in the art based thereon, all of which should be covered within the spirit and scope of this application and the appended claims.
Claims
1. A method of compressing an image, the method comprising: determining an eye gaze position of a user; generating a gaze point map based on the eye gaze position, wherein the gaze point map comprises a first region of the image and a second region of the image; and compressing the first region of the image using a first quality setting and compressing the second region of the image using a second quality setting. Determining the eye gaze position comprises using an eye tracking camera of an augmented reality device.
2. The method of claim 1, wherein, The gaze point map comprises a central region and a peripheral region.
3. The method of claim 1, wherein, The image comprises virtual content generated by an augmented reality device.
4. The method of claim 1, wherein, The image is included in a virtual content video stream.
5. The method of claim 4, wherein, Compressing the first region of the image using the first quality setting comprises compressing all blocks in the first region using the first quality setting.
6. The method of claim 1, wherein, The first quality setting is greater than the second quality setting.
7. The method of claim 1, wherein, The first quality setting is 100%.
8. The method of claim 7, wherein, Post-processing image content in at least one of the first region or the second region.
9. The method of claim 1, further comprising: The compression produces a compressed image, the method further comprising decoding the compressed image using the gaze point map.
10. The method of claim 1, wherein, 11. The method of claim 1, wherein: The first region of the image comprises a plurality of first blocks; The second region of the image comprises a plurality of second blocks; Compressing the first region of the image comprises compressing each of the plurality of first blocks using the first quality setting; and Compressing the second region of the image comprises compressing each of the plurality of second blocks using the second quality setting.
12. The method of claim 1, further comprising: decompressing the first region of the image using the first quality setting; decompressing the second region of the image using the second quality setting; and displaying the image to the user. The second region of the image comprises the first region of the image. The compression produces a compressed image, the method further comprising:
13. The method of claim 1, wherein, decoding the compressed image using the gaze point map to produce a decoded first region and a decoded second region; and 14. The method of claim 13, wherein, reconstructing the image by overlaying the decoded first region on the decoded second region.
15. An augmented reality (AR) system comprising: a wearable device comprising: a frame; a projector coupled with the frame; a display optically coupled with the projector; and an eye tracking system; a memory; and a processor configured to: receive an eye gaze position from the eye tracking system; generate an image; generate a gaze point map based on the eye gaze position, wherein the gaze point map comprises a first region of the image and a second region of the image; and compress the first region of the image using a first quality setting and compress the second region of the image using a second quality setting. The projector comprises one of a set of projectors, the display comprises one of a set of displays, and the eye tracking system comprises a set of eye tracking devices. 16. The AR system of claim 15, wherein, 17. The AR system of claim 16, wherein, Determining the eye gaze position includes using an eye tracking camera of the augmented reality device.
18. The AR system of claim 16, wherein, The gaze point map includes a central region and a peripheral region.
19. The AR system of claim 16, wherein, The image includes virtual content generated by the augmented reality device.
20. The AR system of claim 19, wherein, The image is included in a virtual content video stream.
21. The AR system of claim 16, wherein, Compressing the first region of the image using the first quality setting includes compressing all blocks in the first region using the first quality setting.
22. The AR system of claim 16, wherein, The first quality setting is greater than the second quality setting.
23. The AR system of claim 22, wherein, The first quality setting is 100%.
24. The AR system of claim 16, wherein, The processor is further configured to post-process image content in at least one of the first region or the second region.
25. The AR system of claim 16, wherein, The compression produces a compressed image, and the processor is further configured to decode the compressed image using the gaze point map.
26. The AR system of claim 16, wherein: The first region of the image includes a plurality of first blocks; The second region of the image includes a plurality of second blocks; Compressing the first region of the image includes compressing each of the plurality of first blocks using the first quality setting; and Compressing the second region of the image includes compressing each of the plurality of second blocks using the second quality setting.
27. The AR system of claim 16, wherein, The processor is further configured to: decompress the first region of the image using the first quality setting; decompress the second region of the image using the second quality setting; and display the image to a user. The second region of the image includes the first region of the image.
28. The AR system of claim 16, wherein, The compression produces a compressed image, and the processor is further configured to:
29. The AR system of claim 15, wherein, decode the compressed image using the gaze point map to produce a decoded first region and a decoded second region; and reconstruct the image by overlaying the decoded first region on the decoded second region.
30. A non-transitory computer readable medium comprising program code executable by a processor of a user-wearable device, the program code executable by the processor to: determine an eye gaze position of a user; The gaze point map includes a first region of an image and a second region of the image; generating a gaze point map based on the eye gaze position, wherein and compress the first region of the image using a first quality setting and compress the second region of the image using a second quality setting.